Understanding AI Safety Scores: New Research Questions How We Measure “Harmful Refusal” in AI Models

Photo of author

By Sophia Chen

As artificial intelligence becomes increasingly integrated into everyday life, ensuring these systems behave safely is a top priority. One key safety feature is an AI’s ability to refuse harmful or inappropriate requests, often called “harmful refusal.” A newly published research paper takes a deep dive into how this important trait is measured, revealing that current safety benchmarks may oversimplify complex behaviors. This matters because if we don’t accurately measure AI safety attributes, we risk misunderstanding how well models handle dangerous or policy-violating prompts.

Key Takeaways

  • Current AI safety benchmarks often combine multiple behaviors into a single overall score, which can mask important differences between models.
  • The research focused on “harmful refusal,” the tendency of AI models to reject harmful or policy-violating requests, as a specific safety attribute.
  • Using advanced psychometric analyses, the study found that the main dataset used to measure harmful refusal, HarmBench, does not capture this trait as a single, clear attribute.
  • Aggregating scores across different datasets and behaviors can hide important nuances, suggesting safety scores should be carefully validated before being used to compare models.

To explore how well “harmful refusal” is measured, the researchers examined an influential AI safety benchmark called HELM Safety. Benchmarks like HELM Safety test AI models on multiple datasets designed to evaluate various safety-related behaviors. However, these tests often produce a single overall score, which can make it hard to understand a model’s strengths and weaknesses in specific areas.

The team started by looking at four datasets within HELM Safety that might measure harmful refusal. They found that three of these datasets were “saturated,” meaning they didn’t provide useful, distinct information about refusal behavior. That left HarmBench, a dataset specifically designed to assess harmful refusal, for deeper analysis.

To determine if HarmBench truly measures harmful refusal as a single, unified trait, the researchers applied two psychometric tests. Psychometrics is a field that studies how to measure psychological attributes reliably. The first test, multidimensional item response theory (IRT), checks if a dataset’s questions (or “items”) all relate to one underlying characteristic. The results showed HarmBench items do not align neatly with a single attribute, suggesting harmful refusal is more complex than previously thought.

The second test, differential item functioning (DIF), looks for inconsistencies in how different AI models respond to the same items. The study found that models from different developers with similar refusal abilities sometimes scored differently on individual items. While some of these discrepancies diminished when focusing on specific types of refusal, the pattern indicates that HarmBench’s scores can be influenced by factors beyond a simple refusal trait.

Overall, the paper highlights that HarmBench—and by extension the HELM Safety benchmark’s aggregate scores—combine diverse behaviors into one number. This aggregation can mask important differences between models and obscure the true nature of AI safety behaviors. The authors argue that before using any single score to compare AI models, we must ensure that score accurately reflects one clear, measurable attribute.

These findings have practical implications for AI developers, policymakers, and users who rely on safety benchmarks to evaluate AI systems. More nuanced and validated safety measurements could help identify specific weaknesses in AI behavior, guiding improvements that make these models safer and more trustworthy. Moving forward, the research suggests a need for benchmarks that separate distinct safety traits rather than collapsing them into broad scores, enabling clearer insights into how AI models handle harmful content.

Based on research published on arXiv by Christopher M. Stewart, Preston Botter, Natalie Sarabosing et al..

Editor's note

Editors matched this AI update with related coverage to show where it sits in the broader race over models, regulation and product strategy.

Article briefing

As artificial intelligence becomes increasingly integrated into everyday life, ensuring these systems behave safely is a top...

Story details

  • Author: Sophia Chen
  • Published: October 11, 2026
  • Category: AI

Key developments

  • As artificial intelligence becomes increasingly integrated into everyday life, ensuring these systems behave safely is a top priority.
  • To explore how well “harmful refusal” is measured, the researchers examined an influential AI safety benchmark called HELM Safety.
  • Benchmarks like HELM Safety test AI models on multiple datasets designed to evaluate various safety-related behaviors.

Why this matters

This matters because if we don’t accurately measure AI safety attributes, we risk misunderstanding how well models handle dangerous or policy-violating prompts.

Impact and next steps

More nuanced and validated safety measurements could help identify specific weaknesses in AI behavior, guiding improvements that make these models safer and more trustworthy.

Background

The team started by looking at four datasets within HELM Safety that might measure harmful refusal.

Source

This article is based on source material from arxiv.org.

About the author

Sophia Chen

Sophia Chen covers artificial intelligence and emerging technology. With a background in computer science and a decade of tech journalism, she specialises in AI policy, machine learning applications and the societal impact of automation.

editorial@peacknews.com

Categories AI