Artificial intelligence is increasingly being used to help doctors navigate electronic health records (EHRs), which contain detailed patient information collected over time. However, evaluating how well these AI systems perform is challenging because current testing methods rely on manually created question-and-answer sets that are expensive to produce and quickly become outdated. A newly published research paper introduces an innovative approach to automatically generate and continuously update these evaluation benchmarks, aiming to improve the safety and effectiveness of AI clinical assistants.
Key Takeaways
- Researchers developed a scalable system that automatically creates question-and-answer pairs from patient records spanning multiple doctor visits.
- Nineteen clinicians validated the system, resulting in the Benchmark for Retrieving Information in EHRs (BRIE), a continuously updated test dataset.
- Tests across nine AI language models showed that even advanced systems often miss important clinical details, especially when answers require combining information from multiple documents.
- BRIE enables ongoing, realistic evaluation of AI tools, reflecting diverse clinical reasoning and preventing outdated or leaked test content.
The new framework addresses a major obstacle in deploying AI assistants for healthcare: how to reliably measure their performance in retrieving and synthesizing patient information from complex EHRs. Traditional benchmarks require experts to manually create question-answer pairs, a process that is not only labor-intensive but also becomes obsolete as AI models evolve and medical data changes. To overcome this, the researchers designed an automatic benchmark generator that extracts relevant questions and answers directly from longitudinal clinical notes—records that track a patient’s history over multiple healthcare encounters.
To ensure the quality of this automated approach, the team involved nineteen practicing clinicians to review and validate the generated questions and answers. This collaboration resulted in BRIE, a living benchmark that can be regularly refreshed with new data and variations, reflecting the natural diversity in how different doctors might interpret or prioritize information. This dynamic nature helps guard against “benchmark leakage,” where AI models might overfit to static test sets and appear more accurate than they truly are in real-world settings.
Using BRIE, the researchers evaluated nine state-of-the-art large language models (LLMs) with five different reasoning strategies. The results revealed that even the most advanced AI assistants frequently omit clinically significant details, particularly when questions require synthesizing information across multiple documents or visits—a common scenario in patient care. This insight highlights the ongoing need for rigorous, up-to-date testing to identify AI limitations and guide improvements.
Looking ahead, this scalable and continuously maintainable benchmark framework could play a crucial role in safely integrating AI clinical assistants into healthcare workflows. By providing a realistic and evolving measure of AI performance, BRIE helps ensure these tools support clinicians effectively without missing critical patient information. Future work may expand this approach to cover more diverse patient populations and medical specialties, further enhancing AI’s role in improving healthcare delivery.
Based on research published on arXiv by Jordan L. Cahoon, Chloe O. Stanwyck, Sulaiman Somani et al..
