arXiv:2606. 18043v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets.
By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable.
arXiv:2512.01946v4 Announce Type: replace-cross
Abstract: Robust robotic manipulation requires reliable failure detection and recovery. Although recent Vision-Language Models (VLMs) show promise in r...
By Paul Pacaud, Ricardo Garcia, Shizhe Chen, Cordelia Schmid
arXiv:2607. 08877v1 Announce Type: cross Abstract: Pretrained generative robot policies based on flow matching and diffusion have achieved impressive results across a wide range of manipulation tasks.
By Michael Murray, Daphne Chen, Simran Bagaria, Dean Fortier, Tess Hellebrekers, Galen Mullins, Harshavardhan Gajarla, Oier Mees, Maya Cakmak, Andrey Kolobov
arXiv:2604. 09860v4 Announce Type: replace-cross Abstract: The pursuit of general-purpose robotics has yielded impressive foundation models, yet simulation-based benchmarking remains a bottleneck due to rapid performance saturation and a lack of true generalization testing.
By Jenai Xuning Yang, Rishit Dagli, Alex Zook, Hugo Hadfield, Ankit Goyal, Stan Birchfield, Fabio Ramos, Jonathan Tremblay
arXiv:2608.25757v4 Announce Type: replace-cross
Abstract: Large-scale vision--language--action (VLA) policies have advanced generalist robot control, yet most remain stimulus-to-action black boxes: a...
By Jin Lou, Zhiyuan Jing, Xupeng Wang, Andong Chen, Xingdong Zhu, Yuexuan Li, Yuan Xu, Zhijie Zhu, Yingwei Ji, Wenpeng Nie, Renxing Feng, Liangliang Chen, Ying Chu, Jingyi Li, Jinyan Liu, Zhiqi Song, Jingxuan Zhu, Jidong Zhang, Yufei Liu, Boyang Xing, Lei Jiang, Yan Cui, Hongming Li, Yuchen Zhu
The paper proposes a method to train efficient multi‑task manipulation policies by distilling knowledge from single‑task Conditional Flow Matching (CFM) experts. Instead of training separate models for each task, the authors transfer the experts’ learned velocity fields into a shared policy, combining this distillation signal with the original CFM objective. Experiments on RLBench demonstrate that this approach improves multi‑task performance while keeping the model size fixed, avoiding the need for larger capacity or performance drops seen with naive concatenated training.
By Shreya Deshmukh, Imen Mahdi, Nick Heppert, Abhinav Valada
arXiv:2607. 02092v1 Announce Type: cross Abstract: Flow-matching vision-language-action policies generate robot action chunks through an iterative transport process, creating an opportunity for test-time guidance without retraining the base policy.
By Liuhaichen Yang, Zhuang Jiang, Chenchao Sheng, Zezhi Tang
arXiv:2606. 08881v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have demonstrated strong generalization in robotic manipulation, yet existing evaluations are primarily conducted in simulation or on expensive robotic platforms, leaving their robustness on affordable real-world robots largely unexplored.
By Yi Yu, Xinchuan Qiu
arXiv:2507. 06219v2 Announce Type: replace-cross Abstract: Data scaling has driven remarkable success in foundation models for Natural Language Processing (NLP) and Computer Vision (CV), yet the principles of effective data scaling in robotic manipulation remain insufficiently understood.
By Modi Shi, Li Chen, Jin Chen, Yuxiang Lu, Chiming Liu, Guanghui Ren, Ping Luo, Di Huang, Maoqing Yao, Hongyang Li
arXiv:2607. 14439v1 Announce Type: new Abstract: Generalist robot manipulation policies trained on large, diverse datasets have shown remarkable promise across a wide range of tasks.
By Andrew Liao, Hanchen Cui, Karthik Desingh, Aryan Deshwal
arXiv:2609.35575v2 Announce Type: replace-cross
Abstract: The real-world performance of current vision-language-action models is fundamentally constrained by the limited coverage of expert demonstrat...
By Zhuoyuan Yu, Jiacheng Wang, Tianle Liu, Yihua Ren, Peng Yu, Chen Bai, Ziheng Zhang, Yufei Jia, Jindou Jia, Yuhang Zhang, Xinrui Zhang, Shang Yujing, Yuxiang Chen, Chuhao Zhou, Tiancai Wang, Jianfei Yang