arXiv Machine Learning By Juri Pfammatter, Kaixian Qu, Clemens Schwarke, Victor Klemm, Marco Hutter

Bidirectional Voronoi-biased Exploration Curriculum for Reinforcement Learning

Read the original on arXiv Machine Learning →

The paper introduces Bidirectional Voronoi-biased Exploration Curriculum (BVER), a method that expands start states from the goal and goals from the initial state simultaneously, guiding both toward each other to train a single goal-conditioned policy. Inspired by bidirectional RRT planning, BVER biases exploration toward unexplored task space and steers the two expansions together. Experiments on point-mass mazes, quadrupedal box climbing, and robot-arm ring-on-peg transfer show that BVER learns faster than other reference-free curricula, achieving high success rates and robustness without requiring demonstrations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jun 11

Vision-Language-Action Jump-Starting for Reinforcement Learning Robotic Agents

arXiv:2604. 13733v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) enables high-frequency, closed-loop control for robotic manipulation, but scaling to long-horizon tasks with sparse or imperfect rewards remains difficult due to inefficient exploration and poor credit assignment.

By Angelo Moroncelli, Roberto Zanetti, Marco Maccarini, Loris Roveda
arXiv AI
Aug 28

Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.

By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters