arXiv Machine Learning By Shuang Liu, Yuxuan Bo, Qiuyang Zhao, Caiyue Huang, Xiaorong Chen, Yanguang Liu, Mengnan Du

HARVE: Hacking-Aware Reward-Head Vector Editing for Robust Reward Models

Read the original on arXiv Machine Learning →

arXiv:2606. 03131v1 Announce Type: new Abstract: Reward models are central to large language model (LLM) alignment, but they remain vulnerable to reward hacking.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.