arXiv:2605. 04568v3 Announce Type: replace-cross Abstract: State-of-the-art model-based Reinforcement Learning (RL) approaches either use gradient-free, population-based methods for planning, learned policy networks, or a combination of policy networks and planning.
By Jonathan Spieler, Sven Behnke
arXiv:2607. 21302v1 Announce Type: new Abstract: Behavior prior reinforcement learning (BPRL) has emerged as a promising paradigm to improve sample efficiency in online reinforcement learning (RL) by leveraging policy priors derived from offline demonstrations.
By Gong Gao, Weidong Zhao, Xianhui Liu, Ning Jia
arXiv:2607. 17760v1 Announce Type: cross Abstract: Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations.
By Ziyi Liu, Grace Zhang
Inverse reinforcement learning (IRL) provides a powerful framework for learning from demonstrations. However, real-world tasks often exhibit substantial natural variations (e.
arXiv:2410.14606v3 Announce Type: replace
Abstract: Learning from a stream of experience as it arrives, also known as streaming learning, is a core part of natural learning. However, reliable streami...
By Mohamed Elsayed, Elena Sorina Lupu, Gautham Vasan, A. Rupam Mahmood
arXiv:2609.21108v1 Announce Type: new
Abstract: Deep reinforcement learning (DRL) has achieved strong performance across a wide range of continuous-control problems. These continuous-control policies...
By Sachini Weerasekara, Sagar Kamarthi, Jacqueline Isaacs
The paper introduces Reward Ensemble under Confidence (REC), a probabilistic reward learning framework for preference-based reinforcement learning that models per‑timestep reward uncertainty using an ensemble of distributional reward models. REC incorporates uncertainty into the preference loss and uses model disagreement to drive exploration, achieving 88.4% of shaped‑reward performance on acrobatic quadrotor control versus 55.2% with standard Preference PPO. The authors train policies in simulation and transfer them zero‑shot to real quadrotors, demonstrating complex acrobatic maneuvers learned solely from human preference feedback, and validate REC on a continuous‑control benchmark.
By Colin Merk, Ismail Geles, Jiaxu Xing, Angel Romero, Giorgia Ramponi, Davide Scaramuzza
arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.
By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv:2602. 09580v4 Announce Type: replace-cross Abstract: Real-world fine-tuning of dexterous manipulation policies remains challenging due to limited real-world interaction budgets and highly multimodal action distributions.
By Chenyu Yang, Denis Tarasov, Davide Liconti, Romain Guntz, Hehui Zheng, Robert K. Katzschmann
The paper proposes a reward-based policy that relies only on rewards and actions, enabling zero‑shot transfer between source and target environments with entirely different observation spaces. Experiments on Pointmass, Cartpole, 2D Car Racing, and the Stretch robot in Habitat‑Sim show that the policy can adapt to new visual styles or 3D renderings without additional samples. Additionally, the reward policy can guide the training of an observation‑based policy in the target environment.
By Morgan Byrd, Maks Sorokin, Robert Wright, Sehoon Ha
arXiv:2605. 11020v2 Announce Type: replace-cross Abstract: Inverse reinforcement learning (IRL) is typically formulated as maximizing entropy subject to matching the distribution of expert trajectories.
By Anish Diwan, Davide Tateo, Christopher E. Mower, Haitham Bou-Ammar, Jan Peters, Oleg Arenz
arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.
By Christian Gumbsch, Leonardo Barcellona, Lennard Sch\"unemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves