As artificial intelligence models grow larger and more complex, managing their memory demands becomes a pressing challenge—especially when serving many users at once. A newly published research paper introduces STEPQuant, a novel technique designed to reduce the memory footprint of a key AI component called recurrent states, without significantly sacrificing accuracy. This advancement could help make large AI models more efficient and accessible in real-world applications.
Key Takeaways
- STEPQuant is a new method that compresses recurrent states in AI models by smartly reducing precision where errors matter less, enabling up to 5 times smaller memory use.
- Unlike uniform compression, STEPQuant considers both when (temporal dimension) and where (spatial dimension) quantization errors impact model output, allocating memory resources accordingly.
- Tests on large language models show STEPQuant maintains accuracy close to full-precision (FP32) states using only 6-bit precision, outperforming traditional 8-bit methods at even lower bit rates.
- By integrating STEPQuant with optimized GPU kernels, the approach can cut total serving memory by nearly 70%, improving efficiency for AI services.
To understand STEPQuant, it helps to know what recurrent states are. In some AI models, particularly those using linear attention mechanisms, recurrent states act like a memory that keeps track of information across many steps. Instead of storing ever-growing caches of data, these states hold a fixed-size summary that updates as the model processes new input. However, these states can become a memory bottleneck during concurrent use, such as when multiple users interact with the model simultaneously.
One common strategy to reduce memory is quantization—representing numbers with fewer bits to save space. But simply lowering precision in recurrent states usually harms model accuracy because small errors accumulate over time and across different parts of the state. The researchers behind STEPQuant discovered that not all errors are equally damaging. Temporally, errors in states that persist longer have a bigger impact. Spatially, some parts of the state influence the model output more than others, and the importance varies along both rows and columns of the state matrix.
STEPQuant leverages these insights by allocating precision unevenly: it keeps higher precision where errors would be more harmful and reduces precision where errors have less effect. This spatial-temporal approach involves analyzing the magnitude of the states and how sensitive the model’s output is to errors in each part of the state. The method jointly adjusts scaling factors along both key rows and value columns of the recurrent states to best fit the data distribution and minimize output error.
In experiments conducted on large-scale language models such as Qwen3.8-27B and Kimi-Linear-48B-A3B-Instruct, STEPQuant achieved near full-precision accuracy with only 6 bits per number and outperformed uniform 4-bit quantization. When integrated into the SGLang framework with specialized GPU support, it compressed recurrent states by over five times and cut overall memory use by up to 68.7%, demonstrating significant practical benefits.
This research opens the door to more efficient deployment of large AI models, especially in environments where memory resources are limited or costly. By intelligently deciding when and where to reduce precision, STEPQuant strikes a balance between saving memory and maintaining performance. Future work could explore applying this approach to other model architectures or further optimizing quantization strategies. The researchers have made their code publicly available, inviting the AI community to build on their findings and help bring more scalable AI systems to everyday use.
Based on research published on arXiv by Bingchen Yao, Haobo Xu, Haokun Lin et al..
