As artificial intelligence systems like large language models (LLMs) become more sophisticated, people are asking them complex questions that don’t have clear-cut answers. For example, “What’s the best policy?” or “Which option should I choose?” aren’t questions with objectively correct responses. This poses a challenge for researchers trying to evaluate how well these AI models perform. A newly published research paper suggests that a decades-old approach from economics, known as stated-preference methods, offers a fresh and useful way to assess language models when no ground truth exists.
Key Takeaways
- Evaluating language models on questions without a single correct answer is difficult, but economists have long tackled similar issues using validity frameworks.
- The paper adapts concepts like content validity, construct validity, and reliability from stated-preference economics to the evaluation of AI models.
- Applying these ideas to six language models on an economic survey about water quality revealed clear differences in how well models’ answers aligned with economic theory.
- Passing these validity tests indicates that a model’s responses are internally consistent and coherent, but not necessarily factually correct.
Traditional AI evaluation often relies on comparing model outputs to a known “ground truth” answer. But many real-world questions are subjective or involve complex trade-offs, making such comparisons impossible. To address this, the researchers turned to stated-preference economics, a field that studies how people express their preferences for goods or policies that can’t be directly measured or observed. Economists have developed a set of rigorous validity concepts to judge the quality of survey responses even without knowing the absolute truth.
These concepts include content validity (does the question cover the right topics?), construct validity (does the response relate to the underlying concept it’s supposed to measure?), and criterion validity (does it align with other known measures?). Other important ideas are reliability (are the results consistent?), incentive compatibility (do respondents have motivation to answer truthfully?), and consequentiality (do the responses matter in real decisions?). By translating these terms into the context of language models, the researchers offer a structured way to evaluate AI outputs beyond simple accuracy metrics.
To demonstrate their approach, the team used a published economic survey designed to estimate how much households value improvements in water quality. This survey was given to six different language models, ranging from older to newer versions. Economic theory predicts certain patterns in responses—for example, demand should decrease as price rises, and willingness to pay should increase with income or the scope of the environmental benefit.
The researchers tested whether the models’ answers matched these theoretical expectations. Two older models failed basic validity checks, especially at a household income level of $75,000, while the two newest models passed all tests related to theoretical validity. However, the models still showed differences when assessed for convergent validity, which checks if different methods or measures agree. Importantly, passing these tests doesn’t prove the models are “correct” in an absolute sense, but rather that their answers behave in a coherent, logically consistent manner aligned with economic principles.
This research opens a promising pathway for evaluating AI systems on complex, subjective questions where no definitive answer exists. By borrowing from the toolkit economists have refined over decades, AI evaluators can better understand how well models reason and adhere to expected patterns, rather than simply whether they match a predetermined answer key. Going forward, applying these concepts could improve trust and transparency in AI systems used for policy analysis, decision support, and other areas where nuance and trade-offs are inevitable.
Based on research published on arXiv by Daniel Robert Kling Alexander, Catherine Louise Kling.
