Vision-language models for robots struggle to give consistent feedback when instructions are reworded

Photo of author

By Sophia Chen

As robots become smarter and more versatile, they often rely on artificial intelligence systems called vision-language models (VLMs) to understand instructions and evaluate their own performance. These models combine visual input with natural language commands, allowing robots to learn tasks by associating what they see with what they are told. However, new research reveals a surprising weakness: even slight changes in how instructions are phrased can cause these models to give wildly different evaluations of the same robot behavior. This inconsistency poses a challenge for developing reliable robot learning systems that can understand diverse human instructions.

Key Takeaways

  • Vision-language reward models often assign different success scores to identical robot actions when given paraphrased instructions that mean the same thing.
  • The researchers created ROBORMBENCH, a new benchmark dataset with over 2,300 real robot trajectories and more than 21,000 verified paraphrased instructions to test this problem.
  • Paraphrase-induced inconsistencies are common across many popular proprietary and open-source models, and get worse as the rewordings become more different from the original.
  • Models specifically trained with detailed, trajectory-based feedback show much greater stability and reliability in their reward predictions.

The study focuses on a critical property called “paraphrase invariance.” This means that if two instructions have the same meaning but are worded differently—such as “pick up the red block” versus “grab the crimson cube”—a robust reward model should give the same evaluation of the robot’s attempt to complete the task. If the model’s score changes dramatically just because of different wording, it can mislead the robot during learning, causing it to think it succeeded or failed incorrectly.

To investigate this, the researchers compiled ROBORMBENCH, a comprehensive dataset that includes thousands of real robot task executions paired with a wide variety of paraphrased instructions. These paraphrases cover changes in vocabulary, sentence structure, and even how the goal is described in terms of actions. By testing multiple VLM-based reward models on this dataset, they measured how much the predicted progress scores varied solely due to different wordings of the same instruction.

The results showed that instability is widespread and significant. Even well-known models, including those used in commercial and open-source projects, frequently flipped their judgment of identical robot behaviors from success to failure or vice versa depending on the paraphrase. Interestingly, simply increasing the size of the model or adding explicit reasoning steps did not reliably fix the problem. However, reward models trained with direct supervision grounded in the actual robot trajectories were much more consistent, suggesting that specialized training data is crucial to improve robustness.

These findings highlight a fundamental challenge for using vision-language models as reward functions in robotic learning. If a robot cannot trust the feedback it receives because the evaluation changes with minor wording differences, its learning process can become unreliable or inefficient. Improving paraphrase invariance in VLMs is therefore essential for building robots that can understand and follow diverse human instructions in real-world settings.

Looking ahead, this research opens the door for developing new training strategies and model architectures that emphasize stability across paraphrased commands. The ROBORMBENCH dataset provides a valuable tool for benchmarking progress toward this goal. As robots become more integrated into everyday environments, ensuring they interpret instructions consistently—regardless of phrasing—will be key to safer and more effective human-robot collaboration.

Based on research published on arXiv by Wonje Jeung, Sangyeon Yoon, Hyesoo Hong et al..

Editor's note

This report is framed around the immediate news and the wider implications for regulators, companies and users following the story.

Article briefing

As robots become smarter and more versatile, they often rely on artificial intelligence systems called vision-language models (VLMs) to understand instructions and evaluate...

Story details

  • Author: Sophia Chen
  • Published: September 7, 2026
  • Category: AI

Key developments

  • As robots become smarter and more versatile, they often rely on artificial intelligence systems called vision-language models (VLMs) to understand instructions and evaluate their own performance.
  • However, new research reveals a surprising weakness: even slight changes in how instructions are phrased can cause these models to give wildly different evaluations of the same robot behavior.
  • This inconsistency poses a challenge for developing reliable robot learning systems that can understand diverse human instructions.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI