arXiv Machine Learning

Personalized RewardBench: Evaluating Reward Models with Human Aligned Personalization

arXiv:2604. 07343v2 Announce Type: replace-cross Abstract: Pluralistic alignment has emerged as a critical frontier in the development of Large Language Models (LLMs), with reward models (RMs) serving as a central mechanism for capturing diverse human values.

arXiv Machine Learning
Jul 28

What do Reward Models Memorize?

arXiv:2607. 24484v1 Announce Type: new Abstract: This paper studies what discriminatively trained reward models (RMs) memorize by measuring counterfactual memorization on two human preference datasets.

By Ivo Verhoeven, Pushkar Mishra, Ekaterina Shutova