Training large AI models requires enormous computing power and memory, making it costly and energy-intensive. A key part of this process is the optimizer—a component that helps the model learn by adjusting its parameters. One popular optimizer, called AdamW, keeps track of extra information known as “optimizer states,” which can take up significant storage space. To save memory, researchers have explored quantizing these states—compressing them into fewer bits—but this often introduces errors that hurt training quality. A newly published research paper tackles this problem by redesigning how AdamW’s optimizer states are quantized in 4-bit precision, aiming to reduce errors and improve performance while keeping memory use low.
Key Takeaways
- The study introduces two new quantization methods—ZIP-SR and ZE-EDEN—that round optimizer states in a way that better preserves training accuracy.
- These methods focus on “preconditioner space,” the mathematical space where adaptive updates happen, rather than just the raw optimizer states.
- Experiments on GPT- and Llama-style models (from 130 million to 2.7 billion parameters) show up to a 70% reduction in the accuracy gap compared to previous 4-bit methods.
- Both new approaches perform close to full 32-bit precision optimizers on downstream tasks, demonstrating practical viability for large-scale AI training.
At the heart of this research is the observation that simply minimizing the average error in quantizing optimizer states does not guarantee small errors in the “preconditioner,” a key component that influences how model parameters are updated during training. The preconditioner adjusts learning rates adaptively, making it crucial for maintaining training quality. Traditional quantization methods round values in the “state space,” meaning they decide how to approximate the stored numbers directly. However, this can cause unexpected distortions when these rounded numbers are used in the preconditioner.
To address this, the researchers propose shifting the rounding decision into the “preconditioner space.” This means the quantizer chooses rounding levels based on how the preconditioner will be affected, rather than just the raw stored values. The first method, Zero-Inclusive Preconditioner-space Stochastic Rounding (ZIP-SR), keeps zero as a possible quantized value for the second moment (a statistical measure used by AdamW), and uses probabilistic rounding based on the preconditioner’s values. The second method, Zero-Excluding EDEN calibration (ZE-EDEN), removes zero from the second-moment quantization options but rescales the quantized values to reduce distortion caused by forcing all values to be positive.
Both methods use a specialized 4-bit format called NormalFloat (NF4) for the first moment (another statistical measure in AdamW) and apply targeted stochastic rounding during the final stages of training to further enhance accuracy. The team tested these approaches on a range of transformer models, from smaller 130 million parameter models to very large 2.7 billion parameter models. Results showed that these methods consistently narrowed the gap in validation loss—a key measure of training quality—between 4-bit and full 32-bit precision optimizers. Impressively, the largest improvement reduced this gap by 70% compared to prior 4-bit quantization techniques.
These advancements suggest that carefully redesigning how optimizer states are quantized, by considering the space where adaptive updates happen, can make low-bit optimizers much more reliable. This has important implications for training large AI models more efficiently, potentially lowering the hardware requirements and energy costs without compromising on performance. As AI models continue to grow in size and complexity, such innovations in optimization and quantization will be key to making cutting-edge AI research and applications more accessible and sustainable.
Based on research published on arXiv by Hanyang Li, Shao Tang, Daniel Thomas Braithwaite et al..
