arXiv AI By Peiyu Li, Xiuxiu Tang, Si Chen, Ying Cheng, Ronald Metoyer, Ting Hua, Nitesh V. Chawla

Adaptive Testing for LLM Evaluation: A Psychometric Alternative to Static Benchmarks

Read the original on arXiv AI →

arXiv:2511. 04689v3 Announce Type: replace-cross Abstract: Evaluating large language models (LLMs) typically requires thousands of benchmark items, making the process expensive, slow, and increasingly impractical at scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
6d ago

Efficient Safety Benchmarking via Item Response Theory

The paper demonstrates that Item Response Theory (IRT) can uncover meaningful structure in safety benchmarks for language models, allowing adaptive item selection to approximate full benchmark rankings with Spearman’s ρ > 0.90 while cutting evaluation costs by at least 80% and up to 99.9% on some suites. It also proposes a static method to extract a small, informative subset of items that can be reused across models, achieving 80–99.8% cost savings. These findings show that psychometric techniques can make safety evaluation more efficient without sacrificing ranking accuracy.

By Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri, Diogo Cruz
Hugging Face Trending Papers
Jun 17

LLMs Struggle to Measure What Distinguishes Students of Different Proficiency Levels: A Study of Item Discrimination in Reading Comprehension Assessment

Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While various existing works have explored whether large language models (LLMs) can estimate item difficulty, it remains unclear whether they can capture item discrimination.