arXiv:2609.35823v1 Announce Type: new
Abstract: Vision-language models (VLMs) have made substantial progress in autonomous driving, but their success has primarily been studied in ego-centric scenes....
By Kang Yang, Shuai Liu, Hang Li, Yance Fang, Deying Li, Yongcai Wang
arXiv:2605.25059v4 Announce Type: replace
Abstract: Crucial for autonomous exploration, online 3D occupancy prediction and mapping incrementally construct dense spatial representations on the fly. Em...
By Ruoyu Wang, Yong Liu, Jiahan Li, Sheng Tao, Yuhang Lin, Yukai Ma
arXiv:2608. 07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments.
By Peng Xu, Chengcheng Wang, Shaohua Wan
arXiv:2510.26915v2 Announce Type: replace-cross
Abstract: While heterogeneous teams have typically been designed for well-specified missions with known semantics, generative intelligence, i.e., large...
By Zachary Ravichandran, Fernando Cladera, Ankit Prabhu, Jason Hughes, Carlos Nieto-Granda, Varun Murali, Camillo Taylor, George J. Pappas, Vijay Kumar
arXiv:2606. 27876v1 Announce Type: cross Abstract: Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation.
By Haoyu Zhang, Meng Liu, Qianlong Xiang, Kun Wang, Yaowei Wang, Liqiang Nie
Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets. Despite their strong empirical performance in robotic manipulation, VLAs lack mechanisms to quantify confidence in their predictions and to detect when their actions may be unreliable.
VeriFuse is a bounded arbitration framework that integrates vision‑language models (VLMs) into vehicle‑infrastructure cooperative 3D perception. Each agent first generates independent detections, then VeriFuse creates a unified candidate pool of geometric proposals and cross‑source hypotheses. A frozen VLM selects among three actions—SELECT, REFINE, or REJECT—to resolve ambiguity and produce final 3D detections, achieving strong AP50/AP70 scores on the DAIR‑V2X dataset while keeping vehicle‑side BEV AP50 drop minimal under delay.
By Hongyi Lin, Yiyao Liu, Qi Kang, Heye Huang, Yang Liu, Haris Koutsopoulos, Jinhua Zhao
arXiv:2607. 00283v1 Announce Type: cross Abstract: Autonomous vehicles must safely navigate complex environments where planning-critical agents may be hidden from view.
By Amirhosein Chahe, Tyler Naes, Jovin D'sa, Faizan M. Tariq, Sangjae Bae, Lifeng Zhou, David Isele
arXiv:2606. 18043v1 Announce Type: cross Abstract: Vision-language-action models (VLAs) combine vision-language backbones with expressive generative action heads trained via flow matching on large-scale robotic datasets.
By Ralf R\"omer, Maximilian Seeliger, Saida Liu, Ben Sturgis, Marco Bagatella, Daniel Marta, Andreas Krause, Angela P. Schoellig
arXiv:2608. 04765v1 Announce Type: cross Abstract: Vision-language-action (VLA) models provide a unified paradigm for connecting visual perception, language understanding, and robotic control.
By Houze Xu, Jizhong Li, Ziyi Ye
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
By Yuhong Deng, Yuyao Liu, David Hsu