New Benchmark Challenges AI to Understand Real-World Wearable Health Data

Photo of author

By Sophia Chen

As wearable devices like fitness trackers and smartwatches become increasingly common, they generate vast amounts of health and activity data every day. But can artificial intelligence (AI) systems truly understand and reason about this complex, real-world information? A newly published research paper introduces WearableQA, a benchmark designed to test how well AI models can interpret and analyze long-term wearable health data from real users. This work is important because it moves beyond simple data processing, aiming to evaluate AI’s ability to provide meaningful health insights from continuous, noisy, and varied physiological signals.

Key Takeaways

  • WearableQA includes over 4,000 multiple-choice questions created from daily wearable sensor data, blood test results, and demographic information collected from 200 individuals over up to 500 days.
  • The benchmark tests AI reasoning across 16 question types, focusing on both data computation (analyzing measurements over time) and health reasoning (interpreting physiological meaning).
  • Questions require AI to reason about single signals (like heart rate) as well as integrate multiple signals (such as combining activity and sleep data) for more comprehensive understanding.
  • Evaluation of 14 leading language models showed a wide range of performance, from just under 20% to nearly 73% accuracy, but most models scored below 60%, highlighting the challenge.

The researchers behind WearableQA collected a rich dataset that reflects the real-world conditions of wearable devices, including natural variations between individuals and the presence of device noise—random fluctuations or errors in sensor readings. Instead of relying on synthetic or overly simplified data, they used authentic longitudinal records, meaning continuous daily measurements over many months. This approach ensures that AI models are tested on data that closely resembles what they would encounter outside the lab.

To create the benchmark’s questions, the team developed a “dual-grounding” framework. This means questions are based both on established scientific knowledge from medical literature and on observable patterns validated statistically in the population data. For example, a question might ask an AI to predict how a user’s heart rate changes during sleep based on known physiology and actual recorded trends in the dataset. By combining these two grounding sources, the benchmark captures meaningful, realistic relationships that AI should learn to recognize.

WearableQA’s questions are organized along two main axes. The first axis distinguishes between “data reasoning” (performing calculations or comparisons over time, like finding trends or averages) and “health reasoning” (understanding what physiological changes might mean for a person’s health). The second axis separates questions that focus on single-signal reasoning, which involves analyzing one type of measurement at a time, from cross-signal reasoning, which requires integrating multiple types of data to draw conclusions. This structure allows the benchmark to diagnose specific strengths and weaknesses in AI models.

Testing 14 large language models (LLMs), both proprietary and open-source, the researchers found that while some models performed reasonably well, none solved the benchmark completely. The best models achieved up to 72.9% accuracy, but most scored below 60%, and the lowest were close to random guessing at 10%. This indicates that current AI systems still struggle with the complex task of reasoning over real-world wearable health data.

Looking ahead, WearableQA provides a valuable tool for researchers and developers aiming to improve AI’s understanding of continuous health monitoring data. Better AI reasoning could enhance personalized health insights, early disease detection, and tailored wellness recommendations based on wearable devices. However, the challenge remains significant, and future work will need to focus on improving AI’s ability to handle noisy, multi-dimensional, and longitudinal health data in ways that are medically meaningful.

Based on research published on arXiv by Ji Soo Lee, Xilun Chen, Pierce Chuang et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

As wearable devices like fitness trackers and smartwatches become increasingly common, they generate vast amounts of health and activity data every...

Story details

  • Author: Sophia Chen
  • Published: September 7, 2026
  • Category: AI

Key developments

  • As wearable devices like fitness trackers and smartwatches become increasingly common, they generate vast amounts of health and activity data every day.
  • But can artificial intelligence (AI) systems truly understand and reason about this complex, real-world information?
  • A newly published research paper introduces WearableQA, a benchmark designed to test how well AI models can interpret and analyze long-term wearable health data from real users.

Why this matters

This means questions are based both on established scientific knowledge from medical literature and on observable patterns validated statistically in the population data.

Impact and next steps

To create the benchmark’s questions, the team developed a “dual-grounding” framework.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI