Semifactual Stability Boosts AI Reasoning Accuracy Without Changing Model Weights

Photo of author

By Sophia Chen

Researchers have developed a new technique to make large language models (LLMs) better at reasoning tasks by carefully analyzing how small changes in prompts affect their answers. This approach, called Semifactual Credit-Augmented Policy Optimization (SCAPO), improves the way AI models assign credit to different parts of their generated responses, leading to more accurate and reliable reasoning without needing to retrain the entire model.

Key Takeaways

  • LLMs’ reasoning predictions can be overly sensitive to irrelevant changes in prompts, causing inconsistent answers.
  • Suppressing tokens (words or symbols) that vary widely under small prompt tweaks improves reasoning accuracy without updating model weights.
  • SCAPO, a new training method, uses “semifactual” prompt variations to measure token stability and assign credit more precisely during reinforcement learning.
  • When tested on math benchmarks, SCAPO improved accuracy by over 4 percentage points compared to previous methods and performed best on out-of-distribution tests.

Large language models have shown remarkable ability to solve complex reasoning problems, but their outputs can be surprisingly fragile. Even small, irrelevant changes in the way a question is asked can lead to different answers. This sensitivity suggests that models sometimes rely on spurious cues rather than truly understanding the problem. To tackle this, the researchers introduced the idea of “semifactual” prompt interventions—carefully crafted changes to the prompt that keep the underlying problem and correct answer the same but alter irrelevant details.

By comparing the model’s token-level predictions (the likelihood of each word in the answer) across these semifactual prompts, the team identified which parts of the response were stable and which fluctuated. Tokens that showed high variability—meaning the model was less confident or consistent about them—were considered less reliable. The new method, SCAPO, integrates this stability information into the reinforcement learning process, which is how models improve by receiving feedback on their outputs.

Traditional reinforcement learning approaches like Group Relative Policy Optimization (GRPO) assign the same credit to every token in a response based on the overall outcome. However, this can inadvertently reinforce reliance on unstable or spurious tokens. SCAPO addresses this by adjusting the credit assigned to each token based on its stability score derived from semifactual prompt tests. During early training, tokens that are less stable receive reduced credit, encouraging the model to focus on more reliable reasoning components without penalizing stability alone.

The researchers tested SCAPO on two versions of the Qwen3 language model and evaluated performance on a suite of mathematics reasoning benchmarks from the AIME 2024-2026 dataset and other out-of-distribution tasks. SCAPO outperformed GRPO by 5.63 and 4.17 percentage points on the larger and smaller models respectively. It also achieved the best results on most benchmarks compared to other methods, demonstrating improved reasoning accuracy and better generalization to new problem types.

This work highlights the importance of fine-grained credit assignment in teaching AI models to reason more effectively. By incorporating semifactual stability—essentially measuring how consistent the model’s predictions remain under small, irrelevant prompt changes—SCAPO provides a more nuanced training signal. This could lead to more robust AI systems that are less prone to errors caused by irrelevant input variations.

Looking ahead, the researchers’ approach offers a promising direction for improving the reliability of AI reasoning without the costly process of retraining entire models from scratch. The code for SCAPO is publicly available, inviting further exploration and application across different AI tasks. As language models become increasingly integrated into critical applications, techniques like SCAPO could help ensure their decisions are both accurate and trustworthy.

Based on research published on arXiv by Junshu Pan, Zhizhang Fu, Shulin Huang et al..

Editor's note

This AI briefing pairs the latest development with policy and market context so readers can judge the wider stakes quickly.

Article briefing

Researchers have developed a new technique to make large language models (LLMs) better at reasoning tasks by carefully analyzing how small changes in prompts affect their...

Story details

  • Author: Sophia Chen
  • Published: October 1, 2026
  • Category: AI

Key developments

  • Researchers have developed a new technique to make large language models (LLMs) better at reasoning tasks by carefully analyzing how small changes in prompts affect their answers.
  • Large language models have shown remarkable ability to solve complex reasoning problems, but their outputs can be surprisingly fragile.
  • Even small, irrelevant changes in the way a question is asked can lead to different answers.

Why this matters

Researchers have developed a new technique to make large language models (LLMs) better at reasoning tasks by carefully analyzing how small changes in prompts affect their...

Impact and next steps

This could lead to more robust AI systems that are less prone to errors caused by irrelevant input variations.

Background

The researchers tested SCAPO on two versions of the Qwen3 language model and evaluated performance on a suite of mathematics reasoning benchmarks from the AIME 2024-2026 dataset and other out-of-distribution tasks.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI