arXiv AI By Mohamed Nabail, Leo Cheng, Jingmin Wang, Nicholas Rhinehart

UBP2: Uncertainty-Balanced Preference Planning for Efficient Preference-based Reinforcement Learning

Read the original on arXiv AI →

arXiv:2606. 19328v1 Announce Type: cross Abstract: Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 10

Learning Acrobatic Flight from Preferences

The paper introduces Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for preference-based reinforcement learning that models per‑timestep reward uncertainty using an ensemble of distributional reward models. REC incorporates uncertainty into the preference loss and uses model disagreement to drive exploration, achieving 88.4% of shaped‑reward performance on acrobatic quadrotor control versus 55.2% with standard Preference PPO. The authors train policies in simulation and transfer them zero‑shot to real quadrotors, demonstrating complex acrobatic maneuvers learned solely from human preference feedback, and validate REC on a continuous‑control benchmark.

By Colin Merk, Ismail Geles, Jiaxu Xing, Angel Romero, Giorgia Ramponi, Davide Scaramuzza
arXiv AI
Jun 29

Uncertainty-Aware Reward Discounting for Mitigating Reward Hacking

arXiv:2604. 26360v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncertain about unseen state-action pairs, but the human preference annotations they are trained on are themselves inconsistent, context-dependent, and noisy.

By Disha Singha
arXiv Machine Learning
Aug 24

Smart Exploration in Reinforcement Learning using Bounded Uncertainty Models

The paper introduces BUMEX, a reinforcement learning exploration strategy that leverages a set of prior models containing the true transition kernel and reward function. By optimizing over this model set, the method derives upper and lower bounds on the Q‑function to guide exploration, providing theoretical guarantees of convergence to the optimal policy. When the model set follows a bounded‑parameter MDP structure, the optimization becomes convex, enabling finite‑time convergence under mild assumptions and demonstrating accelerated learning in simulations.

By J. S. van Hulst, W. P. M. H. Heemels, D. J. Antunes