arXiv AI

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.

arXiv AI
Jul 24

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.

By Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, Ranjay Krishna
arXiv Machine Learning
Jul 8

Supervised Reward Inference

arXiv:2502. 18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models.

By Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum
arXiv AI
Aug 28

Residual Reward Models: Leveraging Prior Knowledge for Efficient Preference-based Reinforcement Learning in Robotics

The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.

By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
arXiv AI
Jul 1

Freeform Preference Learning for Robotic Manipulation

arXiv:2606. 32027v1 Announce Type: cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal.

By Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
arXiv Machine Learning
Jun 2

Coherent Off-Policy Improvement of Large Behavior Models with Learned Rewards

arXiv:2606. 02194v1 Announce Type: new Abstract: Distilling expert demonstration data into large generative models using behavioral cloning is a scalable approach to learning capable policies for robotic control, particularly for dexterous manipulation.

By Christian Scherer, Joe Watson, Theo Gruner, Daniel Palenicek, Ingmar Posner, Jan Peters
arXiv Machine Learning
Aug 27

$R^3$: Training Robots to Reason in Natural Language via Reinforcement Learning

The paper introduces $R^3$, a post‑training method that converts vision‑language models into robotic reasoners by first mid‑training on expert reasoning traces and then refining them with single‑step rubric‑based reinforcement learning. $R^3$ enables free‑form language reasoning to guide low‑level manipulation policies, improving exploration, generalization, and performance on long‑horizon tasks in Language Table and simulated bimanual grocery packing benchmarks. The approach outperforms instruction‑only imitation learning baselines and demonstrates that natural language reasoning can serve as a test‑time compute mechanism for steering robotic actions.

By Lehong Wu, Yuxiao Qu, Zheyuan Hu, Ivan Zhang, Limin Wei, Zackory Erickson, Aviral Kumar