arXiv Computation and Language By Sander Land, Daniel M. Bikel

Auditing LLM Benchmarks with Item Response Theory

Read the original on arXiv Computation and Language →

The paper introduces an Item Response Theory (IRT)–based indicator that identifies likely mislabeled items in large language model (LLM) benchmarks with 95% precision among the top 200 examples across seven preference and multiple-choice datasets, using responses from 114 models. It outperforms a supervised classifier and attributes the mislabels to mechanical labeling heuristics, inherited annotation errors, and inherently ambiguous items. The IRT analysis also reveals that reward models tend to specialize in stylistic preference rather than factual knowledge, and pinpoints a frontier reward model that aligns with detected mislabels at 78% accuracy compared to 38% for other models, suggesting benchmark contamination or over‑optimization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Jul 31

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
arXiv Machine Learning
Jun 26

DualEval: Joint Model-Item Calibration for Unified LLM Evaluation

arXiv:2606. 26429v1 Announce Type: new Abstract: Current LLM evaluation relies on two complementary but often disconnected signals: static benchmarks with objective correctness labels and arena-style preference data that better reflect open-ended user interactions.

By Aaron J. Li, Hao Huang, Youngmin Park, Yitong Ma, Wei-Lin Chiang, Li Chen, Cho-Jui Hsieh, Bin Yu, Ion Stoica
arXiv Machine Learning
Jul 28

What do Reward Models Memorize?

arXiv:2607. 24484v1 Announce Type: new Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets.

By Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova
arXiv AI
Aug 20

Self- and Other-Labels Induce Bidirectional Bias in LLM Judges

The study investigates bias in large language model (LLM) judges by having ten LLMs evaluate narrative constraint selections rather than generated text. Results show that self-preference largely disappears under blind evaluation when quality and evaluator severity are controlled, but self- and other-labels alone shift scores bidirectionally when quality is matched. The authors conclude that authorship attribution drives evaluation bias and that open-ended, ground‑truth‑free tasks can effectively study LLM judge behavior.

By Songeun Chae, Min Kim, Donghoon Jung, Seojin Choi, Seohyon Jung
Hugging Face Trending Papers
Jul 27

What do Reward Models Memorize?

This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.