arXiv AI By Disha Singha

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

Read the original on arXiv AI →

arXiv:2604. 26360v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncertain about unseen state-action pairs, but the human preference annotations they are trained on are themselves inconsistent, context-dependent, and noisy.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 9

A Unifying Lens on Reward Uncertainty in RLHF

arXiv:2606. 09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by \emph{reward hacking}, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores without genuine quality gains.

By Ely Hahami, Yoel Zimmermann, Ray Zhou, Jack Benarroch Jedlicki
arXiv Machine Learning
Sep 10

Learning Acrobatic Flight from Preferences

The paper introduces Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for preference-based reinforcement learning that models per‑timestep reward uncertainty using an ensemble of distributional reward models. REC incorporates uncertainty into the preference loss and uses model disagreement to drive exploration, achieving 88.4% of shaped‑reward performance on acrobatic quadrotor control versus 55.2% with standard Preference PPO. The authors train policies in simulation and transfer them zero‑shot to real quadrotors, demonstrating complex acrobatic maneuvers learned solely from human preference feedback, and validate REC on a continuous‑control benchmark.

By Colin Merk, Ismail Geles, Jiaxu Xing, Angel Romero, Giorgia Ramponi, Davide Scaramuzza