Item Response Theory for AI Safety
Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)
Why It Matters
What makes this one worth your time
Understanding and improving the safety of language models is crucial for their reliable deployment in real-world applications, and this paper offers a method to make safety evaluations more efficient and insightful.
IRT is used to enhance the evaluation and auditing of language model safety.
Summary
The paper applies Item Response Theory (IRT) to evaluate the safety of language models across eight benchmarks, analyzing 192 models. It identifies key factors influencing model safety, demonstrates efficient evaluation methods, and suggests IRT as a tool for auditing model behavior.
Key contributions
- Application of IRT to language model safety evaluation.
- Identification of three key factors affecting model safety.
- Demonstration of cost-effective evaluation methods using IRT.
Notable insights
- IRT can reduce evaluation costs by selecting psychometrically significant items.
- IRT can detect naive sandbagging and model changes behind APIs.
Possible limitations
- Not stated in the abstract
Abstract
arXiv:2608.05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks. But aggregated benchmark scores are hard to trust and interpret, because benchmarks duplicate one another, correlate heavily, and models may sandbag when they detect evaluation. To address these issues, we draw on Item Response Theory (IRT), a statistical toolkit for measuring these latents from performance on items with inferred psychometric properties. We fit IRT models to eight safety benchmarks across 192 language models, the largest psychometric analysis of LLM safety evaluations to date, and contribute three results. First, we find that three interpretable factors of refusal strictness, truthfulness, and contextual harm explain most of the variance between models across benchmarks. Second, psychometrically selected items recover full benchmark scores with lower error than random subsets of the same size, and roughly ten adaptively chosen items suffice for several individual benchmarks, cutting evaluation cost by 97-99%. Third, IRT supports audits of individual models, showing that it can be used to detect naive sandbagging and changes of model behind APIs. Overall, we show IRT is a ready-made toolkit for reading, reducing, and auditing safety benchmarks, which we recommend frontier labs and evaluators adopt.