New Transformer Design Lets Language Models Remember More and Think Deeper

Photo of author

By Sophia Chen

Researchers have developed a new approach to improve how transformer-based language models process information, potentially making them more efficient and better at complex reasoning. Traditional transformer models generate text by moving forward layer by layer without revisiting earlier steps or sharing intermediate insights back to previous layers. This limits their ability to remember and refine information during generation. The new method, called the Latent Information Feedback Transformer (LIFT), introduces a way for the model to pass hidden “state” information backward during text generation, helping it keep track of important details and alternative possibilities. This advance could lead to language models that understand context more deeply and perform better on tasks requiring reasoning or following multi-step procedures.

Key Takeaways

  • LIFT adds a feedback mechanism to transformer models, allowing information from deeper layers to be sent back to earlier layers during generation.
  • The model is pretrained using “teacher supervision,” where it learns to predict both the next word and a compact representation of future information derived from an existing pretrained model.
  • Experiments show LIFT outperforms standard transformers on language modeling, reasoning tasks, and procedural tasks, even when using the same number of tokens or computational resources.
  • In tests on tracking state information, a small LIFT model trained with less data beat larger standard transformers trained with much more data, demonstrating efficient learning.

Transformers are the backbone of many modern language models, powering everything from chatbots to translation tools. However, their standard design is feed-forward: information flows in one direction through the network’s layers, and once a layer has passed its output upward, it doesn’t get feedback from deeper layers later on. During text generation, the only way information flows between steps is through the tokens already produced, which limits the model’s ability to reconsider or refine its understanding as it writes.

The research team behind LIFT tackled this by introducing a feedback loop during pretraining. They framed the problem as a “teacher-forced prediction” task, where each input token is paired with a special “state” that contains rich information about what should come next. These states are generated from the next-word predictions of an existing pretrained language model, essentially serving as a guide. The LIFT model then learns to predict both the next token and the next state simultaneously. Because these states are precomputed, training remains efficient and parallelizable across the entire sequence.

At inference time—when the model is generating new text—it feeds back its own predicted states into earlier layers, enabling a flow of information that was previously impossible in standard transformers. This feedback adds only a small computational cost, which becomes even less significant as model size increases.

Testing LIFT on a range of pretrained models from 135 million to 1 billion parameters, the researchers found consistent improvements. The model performed better on language modeling benchmarks, tasks that require reasoning about information, and procedural tasks where following steps correctly is crucial. Notably, in a controlled study focused on tracking hidden state information, a very small LIFT model outperformed same-size standard transformers even when those transformers were trained on eight times more data. This suggests that the feedback mechanism helps the model learn more effectively from less information.

While the research is still at an early stage, LIFT’s approach to incorporating feedback during pretraining could influence the design of future language models. By enabling models to maintain and update a richer internal state throughout generation, it may improve their ability to handle complex instructions, maintain context over longer conversations, and reason more robustly. Further work will be needed to explore how this method scales to even larger models and real-world applications, but the findings open a promising path toward more intelligent and efficient AI language systems.

Based on research published on arXiv by Dor Tirosh, Ido Amos, Mor Geva.

Editor's note

This report is framed around the immediate news and the wider implications for regulators, companies and users following the story.

Article briefing

Researchers have developed a new approach to improve how transformer-based language models process information, potentially making them more efficient and better at complex...

Story details

  • Author: Sophia Chen
  • Published: September 30, 2026
  • Category: AI

Key developments

  • Researchers have developed a new approach to improve how transformer-based language models process information, potentially making them more efficient and better at complex reasoning.
  • This limits their ability to remember and refine information during generation.
  • Transformers are the backbone of many modern language models, powering everything from chatbots to translation tools.

Why this matters

This feedback adds only a small computational cost, which becomes even less significant as model size increases.

Impact and next steps

The research team behind LIFT tackled this by introducing a feedback loop during pretraining.

Background

Traditional transformer models generate text by moving forward layer by layer without revisiting earlier steps or sharing intermediate insights back to previous layers.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI