How a New Approach to AI Models Could Make Them Smarter and More Efficient

Photo of author

By Sophia Chen

Researchers have developed a novel technique to improve a type of artificial intelligence model known as a “Mixture of Experts” (MoE). These models use many specialized components, or “experts,” to handle different parts of a problem, activating only a few experts at a time to save computing power. The new method, called Foil, combines this approach with another technique called “looped transformers,” which reuses the same part of the model multiple times to make better use of its capacity. This fusion promises more efficient and effective AI models that can learn better from data without needing more resources.

Key Takeaways

  • Foil restructures MoE models by flattening expert layers and increasing the number of passes, allowing each routing decision to select from a larger pool of experts.
  • The method unties the attention mechanism across passes, giving each pass its own attention parameters while sharing experts and routing components.
  • Experiments show Foil consistently reduces training loss compared to traditional looped MoE models, indicating better learning efficiency.
  • Foil achieves these improvements without increasing the total number of parameters or computational cost, maintaining or improving accuracy on downstream tasks.

To understand this work, it helps to know a bit about how modern AI models work. Transformers are a popular model architecture that processes data in layers, with each layer building on the previous one. A “looped transformer” takes one such layer and applies it repeatedly, effectively reusing the same parameters multiple times. This can boost the model’s power without adding more parameters, but it has limitations.

Separately, Mixture of Experts models use many different “expert” modules, each specialized in a particular aspect of the task. Instead of activating all experts at once, the model “routes” each input token to only a few experts, saving computation. However, combining looped transformers with MoE designs presents challenges: how do you efficiently reuse experts across loops while maintaining high performance?

The researchers propose Foil as a solution. The key idea is “flattening” the experts—reducing the number of expert layers but increasing the number of experts per layer and the number of passes through the model. This means that at each routing decision, the model chooses from a larger set of experts, enhancing specialization and diversity. Additionally, Foil “unties” the attention mechanism for each pass, allowing each iteration to have its own attention parameters instead of sharing them. Attention is a way for the model to focus on different parts of the input, and untying it helps each pass process information more independently.

Through extensive experiments, the team found that Foil outperforms traditional looped MoE models. At 20 billion training tokens, Foil models achieved lower pretraining loss, meaning they learned more effectively. At 100 billion tokens, the improvements grew even clearer, with the most flattened Foil model reducing loss further while maintaining or improving accuracy on benchmark tasks. The researchers also observed that untying attention led to more confident and balanced routing decisions, which is important for efficient expert usage.

These findings provide practical guidance for designing future looped MoE models: combining more experts per layer with additional passes amplifies benefits, and monitoring routing confidence is a better indicator of healthy expert utilization than simply balancing load. The team has made their code and configurations publicly available, encouraging further exploration and adoption.

Looking ahead, Foil’s approach could help build AI systems that are both more powerful and more resource-efficient, an important goal as models continue to grow in size and complexity. By improving how models use their internal components, researchers can push the boundaries of what AI can learn without proportionally increasing computational demands. This balance will be crucial for making advanced AI accessible and sustainable in practical applications.

Based on research published on arXiv by Shouren Wang, Chuang Ma, Mohsen Hariri et al..

Editor's note

This article focuses on the confirmed update first, then points readers to the competitive and policy context that shapes the beat.

Article briefing

Researchers have developed a novel technique to improve a type of artificial intelligence model known as a "Mixture of Experts"...

Story details

  • Author: Sophia Chen
  • Published: September 29, 2026
  • Category: AI

Key developments

  • Researchers have developed a novel technique to improve a type of artificial intelligence model known as a "Mixture of Experts" (MoE).
  • These models use many specialized components, or "experts," to handle different parts of a problem, activating only a few experts at a time to save computing power.
  • The new method, called Foil, combines this approach with another technique called "looped transformers," which reuses the same part of the model multiple times to make better use of its capacity.

Why this matters

This means that at each routing decision, the model chooses from a larger set of experts, enhancing specialization and diversity.

Impact and next steps

Through extensive experiments, the team found that Foil outperforms traditional looped MoE models.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI