arXiv:2606. 08881v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored.
By Yi Yu, Xinchuan Qiu
arXiv:2607. 29169v1 Announce Type: cross Abstract: Vision-language-action (VLA) policies achieve strong performance in robotic manipulation but remain vulnerable to runtime disturbances that break the temporal alignment among visual observations, robot states, and executed actions.
By Wenda Yu, Tianshi Wang, Fengling Li, Xin Li, Jingjing Li, Lei Zhu
arXiv:2609.21942v1 Announce Type: cross
Abstract: A robot that fails at a task faces the first decision in corrective dialogue: act on its own diagnosis, consult another onboard sensor, or interrupt...
By Eshika Pathak, Leela Krishna
ProTracer is a training‑free framework that uses Vision‑Language Models (VLMs) together with proprioceptive signals to analyze robot manipulation failures. It performs binary failure detection, categorization, explanation generation, and introduces failure onset localization—identifying the earliest moment a robot deviates from a valid trajectory leading to failure. The method leverages proprioceptive dynamics to pinpoint informative action boundaries and converts robot‑state signals into natural‑language descriptions for joint multimodal reasoning, achieving strong performance on both conventional failure diagnosis and the new failure onset localization task.
By Chang Dong, Mehdi Hosseinzadeh, King Hang Wong, Lingqiao Liu, Francois Fraysse, Feras Dayoub, Minh Hoai Nguyen
The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.
By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
INSPECT is a system that learns how a robot should choose its camera view during assembly inspection by observing a smart‑glasses assistant that answers part queries and guides the user. It uses techniques such as Presence‑Invariant TwinSwap for object evidence calibration, claim‑indexed supervision to separate evidence needs from camera changes, and object‑centered calibration to adapt view preferences to robot poses. In experiments on gearbox assemblies and angle‑grinder recordings, INSPECT outperforms other non‑oracle policies, improving view utility and decision accuracy.
By Di Wen, Kailun Yang, Wenhao Guo, Yitian Shi, Junwei Zheng, Yufan Chen, Ruiping Liu, Jiale Wei, Rania Rayyes, Kunyu Peng
arXiv:2604.16683v2 Announce Type: replace-cross
Abstract: Imitation learning has enabled robots to acquire complex visuomotor manipulation skills from demonstrations, but deployment failures remain a...
By Gehan Zheng, Sanjay Seenivasan, Matthew Johnson-Roberson, Weiming Zhi
arXiv:2609.39971v1 Announce Type: cross
Abstract: Vision-language-action (VLA) models can exceed 90% success on in-distribution tasks and withstand nuisance changes that preserve the required action,...
By Hung-Jen Chen, Yu-Hsun Hou, Yan-Hong Chen, Yan-Fu Chen, Binghua Cai, Min Sun, Chun-Yi Lee
arXiv:2608.20784v1 Announce Type: cross
Abstract: Imitation learning for robotics depends on human demonstrations, some of which people may later ask to remove. Retraining without them is the natural...
By Jiazhuo Li, Yu Zhang, Yiming Fei, Kangkang Dong, Xiaojun Zhu, Houde Liu, Jinze Tao
The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.
By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv:2609.08123v1 Announce Type: cross
Abstract: A robot that can be taught a new task from a handful of demonstrations has to work out for itself what it still cannot do, and then ask for exactly t...
By Suyog Khanal, Arun Kumar A V, Santu Rana