AI Legal Assistants Struggle to Spot Missing Details in User Questions, New Study Finds

Photo of author

By Sophia Chen

As artificial intelligence becomes more common in providing legal advice, a new study reveals a critical weakness: AI systems often fail to recognize when a user’s legal question is missing important information. This oversight can lead to inaccurate or misleading answers, which is especially concerning in the complex world of law where small details can dramatically change outcomes. The research introduces a novel benchmark called InsufficiencyBench to test how well AI models detect and respond to incomplete legal queries, highlighting the challenges these systems face in real-world scenarios.

Key Takeaways

  • Most AI legal assistants struggle to identify when a user’s question lacks crucial facts needed for a reliable answer.
  • The study categorizes missing information into eight types across three failure modes, helping clarify why AI models get confused.
  • Among ten leading AI models tested, none performed better than 46% accuracy in spotting missing elements, with many either giving vague hedged answers or making assumptions without warning.
  • No model effectively balances recognizing incomplete queries and confidently answering those that are fully specified.

The research team, led by Samuel J. Vincent and colleagues, developed InsufficiencyBench to address a gap in how AI legal tools are evaluated. Traditional benchmarks assume that all user questions are fully detailed and clear, but in reality, people often omit facts that are legally important—either because they don’t know to include them or because of privacy concerns. This can cause AI systems to jump to conclusions or guess facts, which risks providing incorrect advice.

To tackle this, the researchers first created a detailed taxonomy—a classification system—of the types of missing information that commonly occur in legal questions. These fall into three broad failure modes: “switch,” where a missing fact could change the legal outcome entirely; “gating,” where certain facts are prerequisites for applying a legal rule; and “fatal prerequisite,” where missing information makes it impossible to even start reasoning about the case. They then assembled a dataset of 202 test items, including 58 original queries and 144 deliberately underspecified variants, drawn from six legal areas and 24 US jurisdictions. Practicing attorneys annotated these queries to ensure realistic and meaningful examples.

The team evaluated ten state-of-the-art large language models (LLMs) on this benchmark. They measured the models’ ability to (1) recognize when a question was missing key information, (2) identify specifically what was missing, and (3) avoid giving premature or overly confident answers. The results showed significant shortcomings: the best-performing models could only identify missing elements with moderate accuracy (F2 score up to 0.46), and the median recall—the ability to catch all missing facts—was just 44%. Many models hedged their answers by adding vague disclaimers, while others confidently provided answers based on assumptions that were never made explicit to users.

These findings reveal a critical challenge for AI in legal advice: understanding when it does not have enough information to give a reliable response. The researchers emphasize that no current model successfully balances caution with decisiveness, which is essential for building trust in AI-assisted legal tools.

Looking ahead, this benchmark provides a valuable tool for developers to improve AI legal assistants by making them more sensitive to incomplete queries. Enhancing this capability could lead to safer and more reliable AI advice, helping users avoid costly misunderstandings. As legal AI continues to grow, ensuring these systems recognize their own limitations will be key to their responsible deployment in real-world settings.

Based on research published on arXiv by Samuel J. Vincent, Daniel Calloway, Fangyi Yu et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

As artificial intelligence becomes more common in providing legal advice, a new study reveals a critical weakness: AI systems often fail to recognize when a user's legal...

Story details

  • Author: Sophia Chen
  • Published: August 22, 2026
  • Category: AI

Key developments

  • As artificial intelligence becomes more common in providing legal advice, a new study reveals a critical weakness: AI systems often fail to recognize when a user's legal question is missing important information.
  • This oversight can lead to inaccurate or misleading answers, which is especially concerning in the complex world of law where small details can dramatically change outcomes.
  • The research introduces a novel benchmark called InsufficiencyBench to test how well AI models detect and respond to incomplete legal queries, highlighting the challenges these systems face in real-world scenarios.

Why this matters

This can cause AI systems to jump to conclusions or guess facts, which risks providing incorrect advice.

Impact and next steps

The team evaluated ten state-of-the-art large language models (LLMs) on this benchmark.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI