arXiv:2512. 07019v3 Announce Type: replace-cross Abstract: The proliferation of Large Language Models (LLMs) necessitates valid evaluation methods to provide guidance for both downstream applications and actionable future improvements.
By Zhiyu Xu, Jia Liu, Yixin Wang, Yuqi Gu
arXiv:2601. 02580v2 Announce Type: replace-cross Abstract: Traditional methods for determining assessment item parameters, such as difficulty and discrimination, rely heavily on expensive field testing to collect student performance data for Item Response Theory (IRT) calibration.
By Christopher Ormerod
arXiv:2607. 25257v1 Announce Type: cross Abstract: Item Response Theory (IRT) has recently been proposed as a framework for evaluating large language model (LLM) benchmarks by separating a model's latent ability from the properties of individual benchmark items.
By Juan Francisco, Mandujano Reyes
The paper demonstrates that Item Response Theory (IRT) can uncover meaningful structure in safety benchmarks for language models, allowing adaptive item selection to approximate full benchmark rankings with Spearman’s ρ > 0.90 while cutting evaluation costs by at least 80% and up to 99.9% on some suites. It also proposes a static method to extract a small, informative subset of items that can be reused across models, achieving 80–99.8% cost savings. These findings show that psychometric techniques can make safety evaluation more efficient without sacrificing ranking accuracy.
By Fabio Spagliardi, M\'irian Silva, Ayan Datta, Aiden Zhou, Vamshi Bonagiri, Diogo Cruz
Item discrimination is a fundamental psychometric property of educational assessment, which measures whether an item meaningfully distinguishes students with higher proficiency from students with lower proficiency. While various existing works have explored whether large language models (LLMs) can estimate item difficulty, it remains unclear whether they can capture item discrimination.
arXiv:2601.13885v2 Announce Type: replace-cross
Abstract: Computerized Adaptive Testing (CAT) has proven effective for efficient LLM evaluation on multiple-choice benchmarks, but modern LLM evaluatio...
By Esma Balk{\i}r, Alice Pernthaller, Marco Basaldella, Jos\'e Hern\'andez-Orallo, Nigel Collier