arXiv AI By Jordan Abi Nader, David Lee, Nathaniel Dennler, Andreea Bobu

QuickLAP: Quick Language-Action Preference Learning for Semi-Autonomous Agents

Read the original on arXiv AI →

arXiv:2511. 17855v5 Announce Type: replace Abstract: Robots must learn from both what people do and what they say, but either modality alone is often incomplete: physical corrections are grounded but ambiguous in intent, while language expresses high-level goals but lacks physical grounding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jul 1

Freeform Preference Learning for Robotic Manipulation

arXiv:2606. 32027v1 Announce Type: cross Abstract: Reward design remains a central bottleneck for autonomous robot policy improvement, especially in long-horizon manipulation tasks where sparse success labels provide too little signal and binary preferences collapse many competing notions of quality into one ambiguous signal.

By Marcel Torne, Anubha Mahajan, Abhijnya Bhat, Chelsea Finn
arXiv AI
Jun 2

From Demonstrations to Rewards: Test-Time Prompt Optimization for VLM Reward Models

arXiv:2606. 00083v1 Announce Type: cross Abstract: Reinforcement learning relies on accurate reward functions, which are often hand-crafted or even unavailable in real-world applications, such as robotics.

By Christian Gumbsch, Leonardo Barcellona, Lennard Sch\"unemann, Platon Karageorgis, Andrii Zadaianchuk, Zehao Wang, Sergey Zakharov, Fabien Despinoy, Rahaf Aljundi, Efstratios Gavves
arXiv AI
Jul 24

TOPReward: Token Probabilities as Hidden Zero-Shot Rewards for Robotics

arXiv:2602. 19313v2 Announce Type: replace-cross Abstract: General-purpose robot learning requires dense, instruction-conditioned feedback that can distinguish meaningful task progress from stalled, failed, or partially completed behavior.

By Shirui Chen, Cole Harrison, Ying-Chun Lee, Angela Jin Yang, Zhongzheng Ren, Lillian J. Ratliff, Jiafei Duan, Dieter Fox, Ranjay Krishna
arXiv AI
Sep 15

Rethinking the Implications of Human Feedback for Preference Learning in Human-Robot Collaboration

The paper critiques the standard fixed-rule approach for deriving labels from human feedback in human-robot collaboration, showing that human-provided implication labels often differ and improve reward learning. It introduces IMPLIED, a method that starts with fixed-rule implications but learns to infer and revise accepted/rejected action labels over time, outperforming both the fixed rule and LLM baselines on recorded trajectories and a physical pizza‑making study. As a result, IMPLIED reduces preference‑estimation error and yields robot actions that better align with combined reward objectives.

By Qiping Zhang, Kate Candon, Debasmita Ghose, Marynel V\'azquez