arXiv:2606. 29898v1 Announce Type: cross Abstract: Real-world evaluation is the gold standard for robot policies because it tests them against the physical conditions and deployment challenges they are ultimately designed to handle.
By Haoxu Huang, Tongsam Zheng, Yifan Chen, Jiacheng You, Yang Gao
arXiv:2512.01946v4 Announce Type: replace-cross
Abstract: Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in r...
By Paul Pacaud, Ricardo Garcia, Shizhe Chen, Cordelia Schmid
MAGMA-GEN is an on‑policy data‑generation pipeline that transforms ambiguous failures in hierarchical robotic manipulation into validated recovery supervision. It uses a privileged coach to hypothesize early decision‑level errors and proposes localized corrections, then retains only those candidates that improve downstream progress when re‑executed from the same state. This approach generates supervised examples from the agent’s own failure distribution, enabling improved task success and recovery without requiring per‑step human demonstrations.
By Loan Bernat (LAAS-GEPETTO), Matthieu Grard (LAAS-RAP), Ariane Herbulot (LAAS-RAP), Florent Lamiraux (LAAS-GEPETTO)
arXiv:2608. 07809v1 Announce Type: new Abstract: A world model is only useful for physical AI if it changes what the agent does, and only safe if it declines to do so when it is wrong.
By Yiyao Zhang, Diksha Goel, Hussain Ahmad, Shixun Huang, Jun Shen
arXiv:2607. 03177v1 Announce Type: cross Abstract: Traditional reinforcement learning (RL) for recovery in autonomous systems lacks causal understanding and generalizes poorly to novel failure scenarios.
By Safia Fatima, Kai Olav Ellefsen, Leon Moonen
arXiv:2608. 13719v1 Announce Type: new Abstract: Autonomous systems can fail in rare and heterogeneous ways, making real-world failure discovery difficult under limited testing budgets.
By Anjali Parashar, Rachel Luo, Apoorva Sharma, Sushant Veer, Edward Schmerling, Carson Sobolewski, Mingxin Yu, Chuchu Fan, Marco Pavone
arXiv:2606. 08508v1 Announce Type: cross Abstract: Generative robot policies fail unpredictably at deployment: they hesitate at critical moments, drift off-task, or commit to unrecoverable actions.
By Bingjia Huang, Xiangyu Li, Xiang Wang, Liang Mi, Zixu Hao, Weijun Wang, Hao Wu, Kun Li, Yunxin Liu, Ting Cao
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
arXiv:2607. 01111v1 Announce Type: cross Abstract: Robot policies inevitably encounter failures when deployed in real environments.
By Haoran Hao, Shahram Najam Syed, Jeffrey Ichnowski, Jeff Schneider
arXiv:2606. 08275v1 Announce Type: cross Abstract: When an LLM agent fails -- issues a refund it should not have, calls the wrong tool, leaks data -- existing tooling answers what happened (observability) or whether it passed (evaluation), but not which step caused the failure.
By Jaineet Shah
arXiv:2604.16683v2 Announce Type: replace-cross
Abstract: Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a...
By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
arXiv:2607. 14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks.
By Andrew Liao, Hanchen Cui, Karthik Desingh, Aryan Deshwal