What do Reward Models Memorize?
arXiv:2607. 24484v1 Announce Type: new Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets.
This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets. We show that RMs 1) misallocate memorization to easy, high margin preference pairs, 2) memorize dataset-specific shortcuts (e.
arXiv:2607. 24484v1 Announce Type: new Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets.
Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality. We propose DynaCF, a dynamic reweighting framework for mitigating shortcut learning in reward model training.
arXiv:2606. 09043v1 Announce Type: new Abstract: Reward models trained from pairwise preferences often exploit superficial shortcut cues rather than learning true response quality.
arXiv:2603. 03291v2 Announce Type: replace-cross Abstract: Reward Models (RMs) are crucial for online alignment of language models (LMs) with human preferences.
arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.
arXiv:2509. 22851v4 Announce Type: replace-cross Abstract: Margin-based optimization is fundamental to improving generalization and robustness in classification tasks.
The paper introduces an Item Response Theory (IRT)–based indicator that identifies likely mislabeled items in large language model (LLM) benchmarks with 95% precision among the top 200 examples across seven preference and multiple-choice datasets, using responses from 114 models. It outperforms a supervised classifier and attributes the mislabels to mechanical labeling heuristics, inherited annotation errors, and inherently ambiguous items. The IRT analysis also reveals that reward models tend to specialize in stylistic preference rather than factual knowledge, and pinpoints a frontier reward model that aligns with detected mislabels at 78% accuracy compared to 38% for other models, suggesting benchmark contamination or over‑optimization.
arXiv:2609.14648v1 Announce Type: new Abstract: Aligning multi-turn dialogue agents is usually framed as matching turn-level human preferences, yet direct optimization of long-term outcomes is often...
arXiv:2602. 17658v3 Announce Type: replace-cross Abstract: Reward modeling is central to alignment pipelines such as RLHF, RLAIF, and PPO-based policy optimization, yet its reliability is constrained by limited and heterogeneous human preference data that are expensive to collect at scale.
arXiv:2606. 09124v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards (RLVR) has enabled progress on reasoning-intensive tasks by relying on task-specific verifiers that provide automated correctness signals.
arXiv:2606. 00334v1 Announce Type: cross Abstract: Various language domains have undergone remarkable changes in recent years; these shifts are largely attributed to the advent of Large Language Models and their misalignment with natural language usage.
arXiv:2504.20106v4 Announce Type: replace Abstract: Ensuring that large language models (LLMs) are both helpful and harmless is a critical challenge, as overly strict constraints can lead to excessiv...