New research refines how we measure AI’s problem-solving speed using human task times

Photo of author

By Sophia Chen

As artificial intelligence (AI) systems tackle increasingly complex challenges, understanding how quickly they can solve tasks compared to humans becomes crucial. A newly published study offers a fresh statistical approach to measuring AI capabilities by focusing on “time horizons” — the human time it takes to complete a task that an AI can solve with a 50% success rate. This method helps translate AI performance into more intuitive terms, showing not just what AI can do, but how its abilities scale relative to human effort.

Key Takeaways

  • The study reexamines AI time horizons across 228 tasks and 26 different AI models, improving measurement accuracy using advanced statistical tools.
  • Researchers found that the relationship between human task completion time and AI difficulty is not simply linear, especially for tasks taking between 2 and 30 minutes.
  • The new method offers better estimates of AI time horizons by using splines and item-response theory, which provide a more nuanced “conversion” from human time to AI difficulty.
  • Diagnostic plots introduced in the study help validate the time horizon estimates and should be used alongside benchmarks for more reliable assessment of AI capabilities.

Traditionally, AI researchers have used a metric called the “METR 50% time horizon” to express AI performance: it identifies the amount of time a human would take on a task that an AI can solve half the time. This gives an interpretable scale, linking AI success to human effort. However, previous methods assumed a straightforward, linear relationship between the logarithm of human time and AI difficulty, which may oversimplify the complexity of tasks.

To address this, the researchers applied two statistical techniques: splines and item-response theory (IRT). Splines are flexible mathematical functions that can model curves rather than just straight lines, allowing the relationship between human time and AI difficulty to bend and shift as needed. Item-response theory, often used in educational testing, models how the probability of a correct response depends on both the difficulty of a question and the ability of a test-taker—in this case, adapting it to AI and task difficulty.

By combining these methods, the team developed a model that better captures how AI difficulty relates to human completion time. Notably, the model reveals a “flattened” region for tasks taking between 2 and 30 minutes for humans—meaning that increasing task time within this range doesn’t linearly increase AI difficulty as previously assumed. For example, jumping from a 3-minute task to a 30-minute task is easier for AI than going from 30 minutes to 5 hours, even though both represent a tenfold increase in human time.

Beyond refining the estimates themselves, the researchers introduced diagnostic plots to evaluate how well these time horizon measurements hold up. These visual tools help identify when time horizons provide a valid and consistent picture of AI capabilities and when they might be misleading. The authors recommend that future AI benchmarks using time horizons incorporate such diagnostics, especially as tasks become more complex and time-consuming.

These findings have practical implications for AI development and evaluation. By providing a more accurate and interpretable way to gauge AI problem-solving speed relative to humans, this approach can guide researchers and practitioners in setting realistic expectations and designing better benchmarks. As AI systems continue to evolve, tools like these will be vital for tracking progress and understanding where AI excels or struggles compared to human performance. Future work may extend these methods to even broader task domains and longer time horizons, helping to paint a clearer picture of AI’s growing capabilities.

Based on research published on arXiv by Drew T. Nguyen, William Fithian.

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

As artificial intelligence (AI) systems tackle increasingly complex challenges, understanding how quickly they can solve tasks compared to humans becomes...

Story details

  • Author: Sophia Chen
  • Published: October 9, 2026
  • Category: AI

Key developments

  • As artificial intelligence (AI) systems tackle increasingly complex challenges, understanding how quickly they can solve tasks compared to humans becomes crucial.
  • A newly published study offers a fresh statistical approach to measuring AI capabilities by focusing on “time horizons” — the human time it takes to complete a task that an AI can solve with a 50% success rate.
  • This method helps translate AI performance into more intuitive terms, showing not just what AI can do, but how its abilities scale relative to human effort.

Why this matters

However, previous methods assumed a straightforward, linear relationship between the logarithm of human time and AI difficulty, which may oversimplify the complexity of tasks.

Impact and next steps

By combining these methods, the team developed a model that better captures how AI difficulty relates to human completion time.

Background

Notably, the model reveals a “flattened” region for tasks taking between 2 and 30 minutes for humans—meaning that increasing task time within this range doesn’t linearly increase AI difficulty as previously assumed.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI