This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.
arXiv:2606. 09043v1 Announce Type: new Abstract: Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality.
By Fengyuan Liu, Yongliang Miao, Zirui He, Yanguang Liu, Fei Sun, Mengnan Du
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training.
arXiv:2603. 03291v2 Announce Type: replace-cross Abstract: Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences.
By Daniel Fein, Max Lamparth, Violet Xiang, Mykel J. Kochenderfer, Nick Haber
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
By Qiyao Ma, Dechen Gao, Rui Cai, Boqi Zhao, Hanchu Zhou, Junshan Zhang, Zhe Zhao
The paper introduces an Item Response Theory (IRT)–based indicator that identifies likely mislabeled items in large language model (LLM) benchmarks with 95% precision among the top 200 examples across seven preference and multiple-choice datasets, using responses from 114 models. It outperforms a supervised classifier and attributes the mislabels to mechanical labeling heuristics, inherited annotation errors, and inherently ambiguous items. The IRT analysis also reveals that reward models tend to specialize in stylistic preference rather than factual knowledge, and pinpoints a frontier reward model that aligns with detected mislabels at 78% accuracy compared to 38% for other models, suggesting benchmark contamination or over‑optimization.
By Sander Land, Daniel M. Bikel