arXiv:2606. 32027v1 Announce Type: cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal.
By Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
arXiv:2607. 13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled.
By Chang Liu, Jiawei Zhang, Tao Zhang, Ye Wang, Hongyu Zhou, Qin Jin
arXiv:2502. 18447v2 Announce Type: replace Abstract: Existing approaches to reward inference typically assume that humans provide demonstrations according to specific behavior models.
By Will Schwarzer, Jordan Schneider, Philip S. Thomas, Scott Niekum
arXiv:2511. 17855v5 Announce Type: replace Abstract: Robots must learn from both what people do and what they say, but either modality alone is often incomplete: physical corrections are grounded but ambiguous in intent, while language expresses high-level goals but lacks physical grounding.
By Jordan Abi Nader, David Lee, Nathaniel Dennler, Andreea Bobu
arXiv:2608. 11229v1 Announce Type: new Abstract: Comparative feedback, asking people which of two behaviors they prefer, has become a standard way to align robot and agent behavior with human intent when the reward itself cannot be specified directly.
By Jack Mirenzi, Henny Admoni
arXiv:2608. 08491v1 Announce Type: new Abstract: Reward models are a bottleneck for reinforcement learning in embodied AI.
By Yidong Wang, Yan Zhan, Ziteng Feng, Zhenyu Cui, Ziyi Zhou, Renzhao Liang, Jiaxuan Zhu, Zilei Yang, Yiran Zhao, Zhongkuan Mao, Bo Jia, Hanchu Ni, Chenggang Xie, Biao Liu, Yi Zhang, Yong Dai, Xiaozhu Ju, Wei Ye, Shikun Zhang
The paper presents a unified formalism for proactive robot assistance, organized into three levels, and introduces a framework for unprompted proactive assistance. It demonstrates that offline evaluation overestimates performance and proposes a closed-loop evaluation using an adaptive human model. The authors also present GAP, a method that learns from passive observation to anticipate user goals and act, outperforming prior state-of-the-art methods in closed-loop tests.
By Maithili Patel, Sonia Chernova
The paper introduces Residual Reward Models (RRM) to enhance preference‑based reinforcement learning (PbRL) in robotics. RRMs decompose the true reward into a prior component—such as a heuristic, language‑generated, or IRL‑derived reward—and a learned residual that is trained with human preferences. Experiments on Meta‑World, DM‑Control, and a physical Franka Panda robot show that RRMs markedly improve sample efficiency and accelerate policy learning compared to standard PbRL methods.
By Chenyang Cao, Miguel Rogel-Garc\'ia, Mohamed Nabail, Xueqian Wang, Nicholas Rhinehart
The paper introduces a novel framework that combines vision‑language model (VLM) generated preferences with the Plackett‑Luce (PL) ranking model for reward learning in reinforcement learning. Unlike traditional pairwise Bradley‑Terry approaches, the PL formulation allows listwise rankings of multiple candidates, enabling the use of different ranking sizes (K = 3, 4, 5). Experiments on Meta‑World manipulation tasks show that PL‑based reward models train robotic policies as effectively as, or better than, pairwise, K‑wise, and RL‑VLM‑F baselines, achieving up to an 86% mean final success rate and matching the Oracle baseline on the Drawer Open task.
By Srivalli Katkuri, Maxwell Kawada, Juan Wachs
arXiv:2605. 19294v2 Announce Type: replace-cross Abstract: Vision-Language-Action (VLA) policies increasingly rely on asynchronous inference to hide large-model latency behind ongoing robot motion.
By Yixiang Zhu, Yonghao Chen, Zijie Yang, Yusong Hu, Xinyu Chen
arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.
By Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, Ranjay Krishna
arXiv:2606. 11982v1 Announce Type: new Abstract: Preference-based reinforcement learning (PbRL) learns policies from human trajectory-level comparisons, avoiding explicit reward design and expert demonstrations.
By Aleksandar Taranovic, Onur Celik, Niklas Freymuth, Ge Li, Serge Thilges, Huy Le, Tai Hoang, Rania Rayyes, Gerhard Neumann