arXiv:2606. 19328v1 Announce Type: cross Abstract: Preference-based RL provides an approach to learning reward models from pairwise comparisons of behaviors, bypassing the need for explicit reward design.
By Mohamed Nabail, Leo Cheng, Jingmin Wang, Nicholas Rhinehart
The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.
By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
arXiv:2604. 26360v2 Announce Type: replace-cross Abstract: Reinforcement learning from human feedback (RLHF) systems face a compounding alignment challenge: not only are learned reward models uncertain about unseen state-action pairs, but the human preference annotations they are trained on are themselves inconsistent, context-dependent, and noisy.
By Disha Singha
arXiv:2606. 11982v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations.
By Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv:2606. 06976v1 Announce Type: new Abstract: Large language model (LLM)-based agents often make suboptimal tool-use decisions, including unsupported tool invocation and hallucinated direct responses, which may accumulate errors throughout multi-step interactions.
By Yijin Zhou, Linqian Zeng, Xiaoya Lu, Wenyuan Xie, Dongrui Liu, Junchi Yan, Jing Shao
arXiv:2606.03963v4 Announce Type: replace-cross
Abstract: Deep reinforcement learning enables autonomous robots to learn complex navigation tasks, but still relies heavily on time consuming manual re...
By Roohan Ahmed Khan, Yasheerah Yaqoot, Amir Atef Habel, Muhammad Ahsan Mustafa, Dzmitry Tsetserukou
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
arXiv:2608. 02951v1 Announce Type: cross Abstract: Preference-based reinforcement learning (PbRL) for general stochastic MDPs often requires training a reward model.
By Evan Assmus, Qining Zhang, Lei Ying
arXiv:2607. 11432v1 Announce Type: new Abstract: In this work, we study the reinforcement learning (RL) problem from pairwise trajectory comparisons provided by a human expert.
By Simone Drago, Marco Mussi, Leonardo Bianconi, Alberto Maria Metelli
arXiv:2506. 13741v2 Announce Type: replace-cross Abstract: Preference-based reinforcement learning (PbRL) has emerged as a promising approach for learning behaviors from human feedback without predefined reward functions.
By Brahim Driss, Alex Davey, Riad Akrour
arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.
By Christian Gumbsch, Leonardo Barcellona, Lennard Sch\"unemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves