The paper introduces FailBank, a four‑stage self‑evolving framework that transforms runtime feedback from safety shields into lasting policy improvements for vision‑language‑action (VLA) models. By using a counterfactual correction teacher, outcome‑aware admission, and guarded LoRA updates, FailBank converts useful shield proposals into corrective targets while preserving successful actions as anchors. Experiments on the VLA‑Arena benchmark show that FailBank boosts task success rates by up to 8.5 percentage points and reduces cumulative policy cost by up to 35.6%, outperforming both base policies and traditional runtime shielding.
By Mingyue Cui, Zheyuan Liu, Yihan Zhu, Zheyuan Zhang, Meng Jiang
arXiv:2608. 03231v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) policies promise general robotic manipulation, but their robustness against physical-world attacks remains fragile.
By Jinquan Zhang, Dongfu Yin, Run Yang, Yufeng Yan, Zhen Tian, F. Richard Yu
arXiv:2606. 12299v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models provide a natural language interface to robot control, but the mapping from language to behavior is often brittle and unintuitive: semantically similar instructions can induce drastically different behaviors, while some capabilities may not be elicitable through prompting alone.
By Hyun Joe Jeong, Gokul Swamy, Andrea Bajcsy
arXiv:2609.13231v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models demonstrate strong generalization in robotic manipulation and navigation, but existing fine-tuning methods provid...
By Manan Tayal, Akshay Nambi
The paper introduces FailureSpot, a label‑efficient method for detecting failures at the timestamp level in vision‑language‑action (VLA) policies. It first generates weak supervision from unlabeled VLA action chunks by identifying abnormal patterns, then employs active learning to annotate only the most uncertain trajectories. Experiments on multiple VLA policies demonstrate improved performance for both timestamp‑level and trajectory‑level failure detection.
By Jie Ma, Zongxi Liu, Yi Zhu
arXiv:2607. 04927v1 Announce Type: cross Abstract: World Action Models (WAMs) provide a promising alternative to Vision-Language-Action (VLA) policies by using video-based world modeling as dense supervision for robot action learning.
By Jian Zhu, Jianjun Zhang, Taiyi Su, Tianbin Liu, Zhangyuan Wang, Kai Xie, Zitai Huang, Chong Ma, Youzhang He, Tianjian Wang, Hanyang Wang, Weihao Ding, Yi Xu
arXiv:2609.34792v2 Announce Type: replace
Abstract: Long-horizon manipulation requires robots to remember cues that are no longer in view while responding to moving objects. Yet vision-language-actio...
By Zijian Ye, Chengqi Wei, Wei Huang, Anlin Zheng, Chunyu Zou, Liangyu Wu, Zikang Zhao, Zhenjie Peng, Yushuo Yang, Shuman Zhao, Zhongrui Wang, Xiaojuan Qi
The paper introduces Configured Failure Trapping, a new backdoor attack targeting Vision‑Language‑Action (VLA) models that activates through subtle textual triggers and forces the robot to fail in a specific, controlled manner. It presents a data engine for generating high‑quality target trajectories, an automated evaluation suite, and two benchmarks—Trap‑LIBERO and Trap‑RoboTwin—covering four failure modes. The authors propose TrapVLA, a method that learns trigger‑induced action residuals to steer policies toward the desired failure behavior, demonstrating effectiveness in both simulation and real‑world robotic experiments while maintaining performance on clean data.
By Jun-Hui Liu, Kun-Yu Lin, Yi-Lin Wei, Xu-Han Chen, Yinghao Li, Zhuohao Li, Yuan-Ming Li, Qing Zhang, Xiaoyi Fan, Dongmei Jiang, Yan Li, Wei-Shi Zheng
arXiv:2606. 09572v1 Announce Type: cross Abstract: Vision-language-action models have shown strong promise for robot manipulation, yet raw language is primarily needed to specify task intent rather than to be repeatedly processed during high-frequency low-level execution.
By Jiacheng Li, Yize Guo, Jiabin Guo, Qingchen Liu, Jiahu Qin
CrossSafe proposes embodiment-conditioned safety filtering that uses a Hamilton‑Jacobi reachability value function shared across robots while conditioning on each robot’s morphology and kinematics via a morphology‑aware latent representation. The method performs reachability analysis directly in latent space, enabling a single policy trained on multiple bimanual robot embodiments and manipulation tasks to generalize zero‑shot to a held‑out embodiment and reduce collision rates. Experiments on five embodiments and five tasks demonstrate that training with more embodiments improves generalization.
By Ihab Tabbara, Yuxuan Yang, Hussein Sibai
arXiv:2608.22419v1 Announce Type: cross
Abstract: Query-based Vision-Language-Action (VLA) models offer low-latency inference that is attractive for bimanual robotic manipulation, but we observe that...
By Dongzhou Cheng, Ziang Li, Yixiao Zhou, Haojuan Li, Jinghao Zhang, Lei Lei, Minjing Dong, Jie Gui, Jiaqi Wang
arXiv:2609. 11697v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) and World-Action Models (WAMs) have demonstrated strong capabilities in general-purpose robotic manipulation, yet their generated actions may violate hard physical constraints and therefore be unsafe or infeasible for deployment.
By Jianming Ma, Rongjun Jin, Xiaxi Si, Yang Zhang, Yiheng Li, Yue Gao