arXiv:2609.23997v1 Announce Type: cross
Abstract: Multi-robot collaboration could enable more efficient and scalable solutions to complex robotic tasks, but collaboration under partial observability...
By Dorian Benhamou Goldfajn, Mason Nakamura, Saaduddin Mahmud, Justin Svegliato, Kyle H. Wray, Shlomo Zilberstein
arXiv:2609.36588v1 Announce Type: cross
Abstract: We study reinforcement learning (RL) methods for cooperative multi-agent Vision-Language-Action (VLA) models. This problem is challenging because VLA...
By Ruixiao Xu, Wong Lik Hang Kenny, Zhiqian Liu, Jianing Guo, Hanxiao Li, Kejian Shi, Shuning Zhang, Pu Feng, Yongjia Ma, Yuqing Ma, Kai Chen, Qi Dou, Yaodong Yang, Xianglong Liu, Simin Li
arXiv:2610.02161v1 Announce Type: cross
Abstract: Vision-language models (VLMs) and vision-language-action models (VLAs) have recently driven rapid progress in general-purpose robots, yet most progre...
By Hanchu Zhou, Dechen Gao, Hang Wang, Brendan Lynch, Boqi Zhao, Qiyao Ma, Raman Goyal, Junshan Zhang
arXiv:2607. 22119v1 Announce Type: cross Abstract: Multi-stream robot manipulation policies achieve unparalleled sample efficiency and generalization by modeling actions relative to environmental reference frames.
By Jan Ole von Hartz, Abhinav Valada, Joschka Boedecker
arXiv:2607. 13056v1 Announce Type: cross Abstract: Current vision-language-action (VLA) benchmarks primarily evaluate isolated manipulation skills while leaving human-robot interaction structure largely unmodeled.
By Chang Liu, Jiawei Zhang, Tao Zhang, Ye Wang, Hongyu Zhou, Qin Jin
arXiv:2609.35965v1 Announce Type: cross
Abstract: Vision-and-Language Navigation (VLN) has largely focused on a single agent following a single instruction, yet many real-world applications require t...
By Yunzhe Xu, Zhe Liu
arXiv:2510.26915v2 Announce Type: replace-cross
Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large...
By Zachary Ravichandran, Fernando Cladera, Ankit Prabhu, Jason Hughes, Carlos Nieto-Granda, Varun Murali, Camillo Taylor, George J. Pappas, Vijay Kumar
The paper introduces COMPASS, a decentralized architecture that uses spatial transformers to generate local feedback tokens for large collectives of AI agents in robotics. By aggregating multi‑hop messages, COMPASS enables scalable control of up to 1024 robots, achieving cohesive flocking formations and accurate execution of natural language commands. Experiments show that structured diversity in input commands improves performance and that learned feedback tokens outperform hand‑crafted raw state feedback.
By Frederic Vatnsdal, Roshan Gopal, Romina Garcia Camargo, Vijay Kumar, Alejandro Ribeiro
Vision-Language-Action (VLA) models have attracted growing interest as a scalable approach to robotic manipulation. While these models are effective action predictors, deploying them as robotic agents exposes critical gaps: no mechanism for failure recovery, inconsistent execution over long horizons, and limited robustness to shifts in observations, tasks, or embodiments.
arXiv:2503. 22122v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) have demonstrated remarkable capabilities in robotic planning, particularly for long-horizon tasks that require a holistic understanding of the environment for task decomposition.
By Puzhen Yuan, Angyuan Ma, Yunchao Yao, Huaxiu Yao, Masayoshi Tomizuka, Mingyu Ding
The paper introduces VLAct, a Vision‑Language‑Action model that focuses on representation‑centric continued pre‑training rather than merely scaling robot data. VLAct is trained on diverse, multi‑embodiment robot data and preserves a broad VLM prior while encouraging shared action semantics across embodiments. Experiments across simulation, real‑world, and unseen‑embodiment settings show that VLAct consistently outperforms existing industrial VLA systems, achieving high success rates with only a modest compute budget and open‑source data.
By Senqiao Yang, Chengyao Wang, Yuxin Chen, Zixuan Wang, Longxiang Tang, Haokun Gui, Jinhui Ye, Changsheng Lu, Xiaoyang Wu, Mingkang Zhu, Pengguang Chen, Shu Liu, Zhuotao Tian, Hengshuang Zhao, Bei Yu, Jiaya Jia
arXiv:2607. 09776v1 Announce Type: cross Abstract: When adapting Vision Language Action (VLA) models to downstream tasks, multiple rounds of post training are required because a single round of data cannot resolve all issues, making continuous iterations necessary to progressively address the weaknesses exposed in previous rounds.
By Shaopeng Zhai, Qi Zhang, Tianyi Zhang, Haoran Zhang, Fuxian Huang, Zhanhui Lin, Zijun Xu