arXiv Machine Learning By Yikai Wang, Shang Liu, Jose Blanchet

Wasserstein Distributionally Robust Regret Optimization for Reinforcement Learning from Human Feedback

Read the original on arXiv Machine Learning →

arXiv:2605. 00155v3 Announce Type: replace Abstract: Reinforcement learning from human feedback (RLHF) is a central post-training tool for aligning large language models, but its training reward is only a learned proxy for true human utility.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jun 30

Wasserstein Distributionally Robust Regret Optimization

arXiv:2504. 10796v4 Announce Type: replace-cross Abstract: Distributionally robust optimization (DRO) is widely used for decision-making under uncertainty, but its adversarial focus on worst-case loss can lead to overly conservative policies.

By Lukas-Benedikt Fiechtner, Jose Blanchet
arXiv AI
Jun 9

A Unifying Lens on Reward Uncertainty in RLHF

arXiv:2606. 09073v1 Announce Type: cross Abstract: Reinforcement learning from human feedback (RLHF) is bottlenecked by \emph{reward hacking}, where the policy exploits errors in a proxy reward model (RM) and produces high RM scores without genuine quality gains.

By Ely Hahami, Yoel Zimmermann, Ray Zhou, Jack Benarroch Jedlicki