arXiv AI By Joshua Fonseca Rivera (Independent), Neil Shah (Independent), David Demitri Africa (UK AI Security Institute), Konstantinos Voudouris (UK AI Security Institute)

Item Response Theory for AI Safety

Read the original on arXiv AI →

arXiv:2608. 05086v1 Announce Type: new Abstract: Language models differ in how safely they behave and these differences are measured by safety benchmarks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Efficient Safety Benchmarking via Item Response Theory

The paper demonstrates that Item Response Theory (IRT) can uncover meaningful structure in safety benchmarks for language models, allowing adaptive item selection to approximate full benchmark rankings with Spearman’s ρ > 0.90 while cutting evaluation costs by at least 80% and up to 99.9% on some suites. It also proposes a static method to extract a small, informative subset of items that can be reused across models, achieving 80–99.8% cost savings. These findings show that psychometric techniques can make safety evaluation more efficient without sacrificing ranking accuracy.

By Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri, Diogo Cruz
arXiv AI
Aug 19

Benchmarking the Benchmarks: Evaluating Automated Safety Benchmarks for Small Language Models

The paper investigates whether existing AI safety benchmarks, designed for large language models, are suitable for evaluating small language models (SLMs). By testing five benchmark suites on 26 open‑source SLMs with a unified scoring rubric, the authors find that ambiguous judgments dominate, especially for complex prompts and certain architectures. This ambiguity, linked to factors like lexical density and output perplexity, undermines the reliability of aggregate leaderboards and reveals a confound between model capability and perceived safety.

By Nyamtulla Shaik, Fengjun Li, Bo Luo