New Benchmark Challenges AI Agents to Master Multi-Device Workflows Across Platforms

Photo of author

By Sophia Chen

As people increasingly juggle tasks across smartphones, laptops, and desktops, the ability for AI assistants to navigate and coordinate actions across multiple devices is becoming crucial. However, most current AI systems designed to interact with graphical user interfaces (GUIs) are tested only on single devices performing isolated tasks. A new research paper introduces JarvisGUI, a benchmark designed to push AI agents to handle complex workflows that span different devices and operating systems—reflecting the realities of everyday digital life more accurately.

Key Takeaways

  • JarvisGUI is a novel benchmark that tests AI agents on cross-device GUI tasks involving Android, Windows, and Ubuntu platforms.
  • The benchmark uses a lightweight type system to automatically create multi-step workflows requiring agents to transfer information and maintain shared states across devices.
  • Current state-of-the-art open-source GUI agents struggle with the challenges of cross-platform context understanding and managing dependencies over long, multi-device tasks.
  • This exposes a significant gap in existing AI capabilities that previous single-device benchmarks have overlooked.

The researchers behind JarvisGUI argue that real-world use of AI assistants often involves workflows where intermediate results need to be passed seamlessly between devices—for example, starting a task on a smartphone and finishing it on a desktop computer. Existing evaluation methods typically focus on single-device environments with fixed tasks, which do not capture the complexities of these dynamic, multi-device workflows.

To address this, JarvisGUI formulates GUI tasks as input-output transformations governed by a lightweight type system. In simpler terms, this system classifies the kinds of data or information being passed between steps, enabling the automatic assembly of complex workflows that require coordination across different platforms. By running AI agents in virtual environments that mimic Android smartphones, Windows PCs, and Ubuntu machines, the benchmark tests whether these agents can maintain shared states, transfer intermediate results correctly, and reason about the context on each device.

“Cross-device workflows introduce challenges like state-transfer awareness—knowing what information to carry over between devices—and cross-platform contextual reasoning, which means understanding how tasks relate across different operating systems,” explained the authors. JarvisGUI also evaluates how well agents manage long-horizon dependencies, meaning the ability to keep track of multiple steps over time without losing track of earlier context.

When tested on this benchmark, leading open-source GUI agents showed significant difficulty in handling these requirements. This suggests that despite progress in GUI automation and AI interaction, current systems are not yet ready to support the fluid, multi-device workflows common in everyday user scenarios.

The introduction of JarvisGUI highlights an important direction for future AI research: developing agents that can operate seamlessly across diverse hardware and software environments. This could have practical implications for improving virtual assistants, automated workflows, and user productivity tools that rely on AI. The benchmark offers a unified framework for researchers to measure progress in this area and to identify the specific challenges that need to be overcome.

As digital ecosystems continue to grow more interconnected, tools like JarvisGUI will be essential to ensure AI agents evolve beyond isolated tasks and single-device interactions—paving the way for more intelligent, adaptable assistants capable of managing our increasingly complex digital lives.

Based on research published on arXiv by Zixiang Chen, Yuheng Lu, Zihao Cheng et al..

Editor's note

This report is framed around the immediate news and the wider implications for regulators, companies and users following the story.

Article briefing

As people increasingly juggle tasks across smartphones, laptops, and desktops, the ability for AI assistants to navigate and coordinate actions across multiple devices is...

Story details

  • Author: Sophia Chen
  • Published: September 10, 2026
  • Category: AI

Key developments

  • As people increasingly juggle tasks across smartphones, laptops, and desktops, the ability for AI assistants to navigate and coordinate actions across multiple devices is becoming crucial.
  • However, most current AI systems designed to interact with graphical user interfaces (GUIs) are tested only on single devices performing isolated tasks.
  • Existing evaluation methods typically focus on single-device environments with fixed tasks, which do not capture the complexities of these dynamic, multi-device workflows.

Why this matters

This could have practical implications for improving virtual assistants, automated workflows, and user productivity tools that rely on AI.

Background

JarvisGUI also evaluates how well agents manage long-horizon dependencies, meaning the ability to keep track of multiple steps over time without losing track of earlier context.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI