arXiv AI By Zhibin Duan, Guowei Rong, Zhuo Li, Bo Chen, Mingyuan Zhou, Dandan Guo

Mitigating Reward Hacking in RLHF via Bayesian Non-negative Reward Modeling

Read the original on arXiv AI →

arXiv:2602. 10623v2 Announce Type: replace-cross Abstract: Reward models learned from human preferences are central to aligning large language models (LLMs) via reinforcement learning from human feedback, yet they are often vulnerable to reward hacking due to noisy annotations and systematic biases such as response length or style.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.