arXiv:2608.29537v1 Announce Type: cross
Abstract: Frozen vision-language-action (VLA) policies offer broad manipulation skills but execute open-loop action chunks without tracking task progress, so t...
By Hongbo Gao, Zeyu Ni, Xin Wen, Siyu Xu, Ruifeng Li
EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.
By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments.
The paper introduces 2AM, a system that keeps task memory solely within a multimodal Agent while using a single RGB‑based, stateless Action Model to execute motions. By compiling interaction history into subtask language and optional 2D grasp/place/move hints, the Agent steers the Action Model, which is trained to tolerate imperfect guidance through dropout, noise, and jitter. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion without depth, geometry, or planners, vastly outperforming the best baseline.
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang
The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.
By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.
By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
The paper introduces Hide-and-Seek, a framework for detecting failures in Vision‑Language‑Action (VLA) models during robot execution. It treats failure detection as a coarsely supervised learning problem, using inter‑trajectory and intra‑trajectory contrastive objectives to localize failure‑indicative actions without step‑level annotations. Experiments on LIBERO, VLABench, and a real‑world robotic platform show that Hide‑and‑Seek achieves state‑of‑the‑art multi‑task failure detection performance across several VLA policies.
By Seongheon Park, Wendi Li, Changdae Oh, Samuel Yeh, Zsolt Kira, Michael Hagenow, Sharon Li
arXiv:2607. 20345v1 Announce Type: cross Abstract: Closing the gap between benchmark performance and reliable real-world operation remains a central challenge for Vision-Language-Action (VLA) humanoid robots, which must handle execution errors, distribution shifts, and environmental variability.
By Roger Sala Sis\'o, Tiago Silv\'erio, Jakob Sand, Tran Nguyen Le
arXiv:2606. 17011v1 Announce Type: cross Abstract: Human interventions provide crucial corrective signals for post-training Vision-Language-Action (VLA) models.
By Wei Xiao, Weiliang Tang, Yuying Ge, Hui Zhou, Yao Mu, Li Zhang, Yixiao Ge
Embodied planning increasingly relies on vision-language models (VLMs) to translate instructions and visual observations into executable action sequences. However, fluent plans are not always executab...
arXiv:2512.01946v4 Announce Type: replace-cross
Abstract: Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in r...
By Paul Pacaud, Ricardo Garcia, Shizhe Chen, Cordelia Schmid