Researchers have developed a new approach to improve how artificial intelligence (AI) models learn to solve complex problems, particularly in mathematical reasoning. The study focuses on reinforcement learning with verifiable rewards (RLVR), a technique where AI models are trained to generate correct answers and are rewarded when their solutions can be verified. While RLVR holds promise for discovering novel reasoning strategies, encouraging models to explore new ideas without losing accuracy has been challenging. This new research introduces a method called Exploration-Distillation (ExpDis) that separates the process of exploring novel solutions from optimizing for correctness, leading to better performance without degrading the model’s quality.
Key Takeaways
- Exploration-Distillation (ExpDis) separates the discovery of new ideas from the fine-tuning of correct solutions, improving AI reasoning.
- By training “explorer” models with a novelty bonus and then distilling their best outputs into a “student” model without novelty incentives, ExpDis avoids quality degradation.
- ExpDis outperforms previous methods like DAPO on seven mathematical reasoning benchmarks across different AI model families.
- The approach leads to more diverse correct solutions, enhancing the model’s ability to generate multiple valid answers.
Reinforcement learning with verifiable rewards (RLVR) is a training method where AI models receive feedback based on how well their outputs can be verified as correct. This technique is valuable for tasks like mathematical reasoning, where solutions can be checked for accuracy. However, simply encouraging models to explore novel or creative answers by adding a “novelty bonus” to their rewards often backfires, causing the AI to produce lower-quality or incorrect solutions. This happens because the narrow focus on verifiable rewards doesn’t cover the full range of knowledge and behaviors the model has learned, making it hard to recover from mistakes introduced during exploration.
The new method, Exploration-Distillation (ExpDis), tackles this problem by splitting the learning process into two distinct phases. First, “explorer” policies are trained with a novelty bonus to encourage them to try out new ideas and generate diverse solutions. These explorer models produce many candidate answers, which are then carefully filtered to select those that are both novel and correct. Next, the filtered solutions are used to train a separate “student” policy that learns to produce high-quality answers without any novelty incentives. By alternating between these exploration and optimization phases over multiple rounds, ExpDis allows the AI to safely explore new strategies while maintaining or improving overall solution quality.
This approach was tested on seven different mathematical reasoning benchmarks using two different AI model families. Across these tests, ExpDis consistently outperformed a previous method known as DAPO, achieving better results within the same amount of training time. Notably, ExpDis also improved “pass@$k$” scaling, a measure indicating that the model generates a greater number of diverse correct solutions rather than repeating the same answers. This diversity is important in fields like mathematics and programming, where multiple valid approaches may exist for solving a problem.
By decoupling exploration from optimization, this research offers a promising way to encourage AI models to discover new reasoning strategies without sacrificing accuracy or reliability. While the current work focuses on mathematical reasoning, the underlying principles could be applied to other domains where AI must balance creativity and correctness. Future research may explore how ExpDis can be integrated with larger language models or adapted for real-world applications such as scientific discovery, automated programming, or complex decision-making. As AI systems continue to evolve, methods like ExpDis could help unlock more advanced reasoning capabilities while maintaining trustworthiness and precision.
Based on research published on arXiv by Saif Punjwani, Micah Goldblum.
