A New Approach Helps AI Understand Human Preferences Without Complex Calculations

Photo of author

By Sophia Chen

As artificial intelligence grows more powerful, getting AI systems—especially large language models (LLMs)—to align with human preferences becomes increasingly important. A newly published research paper introduces a fresh method that could make this alignment more efficient and reliable. Instead of relying on traditional techniques that directly optimize how well models match human choices, the researchers propose a novel approach called Comparison-based Preference Optimization (ComPO). This method uses simpler comparison signals to guide AI learning, potentially reducing computational demands while improving performance.

Key Takeaways

  • ComPO is a new “zeroth-order” method that aligns language models with human preferences using comparison oracles, which assess which output is better without needing detailed gradient calculations.
  • The method provides theoretical guarantees of convergence under certain conditions, meaning it reliably improves model alignment over time.
  • ComPO works both offline—using pre-collected preference data—and online, incorporating new, unlabeled model outputs to refine alignment dynamically.
  • Experiments on several leading AI models, including Mistral, Llama, and Gemma series, show ComPO outperforms existing direct alignment techniques, especially in controlling output length and preference accuracy.

To understand why this matters, it helps to know how AI models are usually trained to follow human preferences. Traditionally, researchers collect pairs of AI-generated outputs and label which one humans prefer. Then, the model is fine-tuned to increase the likelihood of preferred outputs using gradient-based optimization—a way to adjust model parameters by calculating precise derivatives of a loss function. However, this approach can struggle when the differences between preferred and non-preferred outputs are subtle, a challenge known as “likelihood displacement.”

ComPO takes a different tack by treating preference alignment as a “zeroth-order” optimization problem. In optimization jargon, zeroth-order methods don’t require gradients (derivatives) to find better solutions—instead, they rely only on comparisons or function evaluations. Here, the “comparison oracle” is a mechanism that can tell which of two outputs is better but doesn’t provide detailed numerical feedback. Using these comparisons, ComPO extracts directional information that guides the model towards better preferences without explicitly calculating gradients of a preference loss.

The researchers developed both offline and online versions of ComPO. The offline version uses a fixed set of preference pairs and guarantees that the model’s performance will improve under assumptions like smoothness (small changes in parameters lead to small changes in outcomes) and gradient sparsity (only a few parameters strongly affect preferences). The online version extends this by incorporating new, unlabeled model outputs, controlling the model’s behavior relative to a reference policy through a technique called reverse-KL divergence—a statistical measure of how one probability distribution differs from another. This allows the model to learn continuously and adaptively from comparisons while maintaining stable and controlled updates.

Testing ComPO on state-of-the-art language models like Mistral, Llama, and multiple versions of Gemma and Qwen3, the researchers found consistent improvements over standard alignment methods. Notably, ComPO helped models better manage the length of their responses—a common challenge in AI-generated text—and increased “win rates” in preference comparisons. Additional diagnostics at the pair level suggest that ComPO effectively mitigates the likelihood displacement problem, leading to more reliable alignment with human choices.

While this research is still early-stage, the implications are promising. By reducing reliance on gradient calculations and enabling more flexible, comparison-based learning, ComPO could make aligning AI systems with human values more efficient and scalable. This is particularly relevant as language models grow larger and more complex, where computational costs and alignment challenges increase. Future work may explore applying ComPO to other AI domains beyond language modeling and integrating it with existing training pipelines to enhance AI safety and usability.

Based on research published on arXiv by Peter Chen, Xi Chen, Wotao Yin et al..

Editor's note

This report is framed around the immediate news and the wider implications for regulators, companies and users following the story.

Article briefing

As artificial intelligence grows more powerful, getting AI systems—especially large language models (LLMs)—to align with human preferences becomes increasingly...

Story details

  • Author: Sophia Chen
  • Published: September 17, 2026
  • Category: AI

Key developments

  • As artificial intelligence grows more powerful, getting AI systems—especially large language models (LLMs)—to align with human preferences becomes increasingly important.
  • Instead of relying on traditional techniques that directly optimize how well models match human choices, the researchers propose a novel approach called Comparison-based Preference Optimization (ComPO).
  • This method uses simpler comparison signals to guide AI learning, potentially reducing computational demands while improving performance.

Why this matters

A newly published research paper introduces a fresh method that could make this alignment more efficient and reliable.

Impact and next steps

By reducing reliance on gradient calculations and enabling more flexible, comparison-based learning, ComPO could make aligning AI systems with human values more efficient and scalable.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI