WOVEN advances AI’s understanding of visual changes to boost multimodal language models

Photo of author

By Sophia Chen

Researchers have developed a new approach to help artificial intelligence better understand how the visual world changes over time—a crucial step toward improving AI’s reasoning about space, physical interactions, and sequences of events. The study, recently published on arXiv, introduces WOVEN, a large-scale training dataset and benchmark designed to teach multimodal large language models (MLLMs) to reason about visual transitions. This work addresses persistent challenges in AI systems that combine language and vision, such as understanding how objects move, change, or interact in complex scenes.

Key Takeaways

  • Current leading multimodal AI models struggle with reasoning about visual transitions—how scenes change due to actions—falling well short of human performance.
  • WOVEN is a new dataset containing over 36,000 examples of visual transitions, organized by types of scenes, actions, and reasoning tasks, created using video-based generative models.
  • Training models on just small subsets of WOVEN data significantly improves their performance across 22 out of 26 external visual reasoning benchmarks, in some cases by over 27 percentage points.
  • The research provides a practical “training recipe” emphasizing the importance of focusing on reasoning operations and larger visual changes to build more robust visual world understanding in AI.

Multimodal large language models are AI systems designed to process and generate information across different types of data, such as text and images. While these models have made impressive strides in understanding language and recognizing objects in images, they often falter when asked to reason about how visual scenes evolve—like predicting what happens next after an action or understanding the physical consequences of interactions. The researchers behind WOVEN hypothesized that many of these challenges stem from a shared underlying difficulty with “visual transition reasoning,” which involves understanding the changes in a scene over time due to various actions.

To test this idea, the team created WOVEN, a comprehensive dataset and benchmark that systematically organizes examples of visual changes by scene type (e.g., kitchen, playground), action type (e.g., moving, stacking), and reasoning type (e.g., spatial, temporal). The data was generated using advanced video-pretrained generative models, providing realistic and diverse “rollouts” showing how scenes change step-by-step. This structure allowed the researchers to clearly evaluate and compare how well different state-of-the-art MLLMs, including models like GPT-5.4 and Qwen3-VL, perform on this kind of reasoning.

The findings revealed a consistent pattern: even the most advanced models struggled with visual transition reasoning, performing far below human levels. These difficulties were widespread across different model architectures and did not improve simply by scaling up model size. However, when the researchers trained these models on WOVEN, even small subsets of the data led to large gains. Remarkably, models trained on just about 2,000 examples from WOVEN improved their accuracy on a wide range of external benchmarks, demonstrating that visual transition reasoning is a transferable skill that benefits many tasks.

In addition to the dataset itself, the study proposes a “training recipe” to maximize learning efficiency. Instead of focusing training on specific scenes or action types, it’s more effective to organize supervision by the type of reasoning operation being taught. Moreover, emphasizing examples with larger visual changes helps models develop more robust understanding. This approach was validated on benchmarks the models had not seen before, confirming its practical value for future training regimes.

By establishing visual transition reasoning as a reusable foundation, WOVEN opens the door to more systematic and scalable training of multimodal AI systems. This progress could enhance applications ranging from robotics and autonomous vehicles that need to predict physical interactions, to assistive technologies that interpret dynamic environments. While challenges remain, the research marks a meaningful step toward AI models that better understand the complex, changing visual world around us.

Based on research published on arXiv by Zheyu Fan, Yue Zhang, Mingkai Deng et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

Researchers have developed a new approach to help artificial intelligence better understand how the visual world changes over time—a crucial step toward improving AI’s...

Story details

  • Author: Sophia Chen
  • Published: October 10, 2026
  • Category: AI

Key developments

  • The study, recently published on arXiv, introduces WOVEN, a large-scale training dataset and benchmark designed to teach multimodal large language models (MLLMs) to reason about visual transitions.
  • This work addresses persistent challenges in AI systems that combine language and vision, such as understanding how objects move, change, or interact in complex scenes.
  • The data was generated using advanced video-pretrained generative models, providing realistic and diverse “rollouts” showing how scenes change step-by-step.

Why this matters

This progress could enhance applications ranging from robotics and autonomous vehicles that need to predict physical interactions, to assistive technologies that interpret dynamic environments.

Background

This approach was validated on benchmarks the models had not seen before, confirming its practical value for future training regimes.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI