arXiv Machine Learning By Yikai Wang, Shang Liu, Jose Blanchet

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

Read the original on arXiv Machine Learning →

arXiv:2605. 00155v3 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

Wasserstein Distributionally Robust Regret Optimization

arXiv:2504. 10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies.

By Lukas-Benedikt Fiechtner, Jose Blanchet
arXiv AI
Jun 9

A Unifying Lens on Reward Uncertainty in RLHF

arXiv:2606. 09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by \emph{reward hacking}, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores without genuine quality gains.

By Ely Hahami, Yoel Zimmermann, Ray Zhou, Jack Benarroch Jedlicki