Air-Ground Collaborative Vision-and-Language Navigation (AGC-VLN) pairs an unmanned aerial vehicle (UAV) with a global bird’s‑eye view and an unmanned ground vehicle (UGV) with a local first‑person view, creating a shared bird’s‑eye map that enables collaboration. The training‑free baseline decomposes navigation into VLM‑based semantic reasoning and deterministic geometric execution, allowing the UAV to render the UGV’s pose and target markers while the UGV plans a road‑following path using the shared map. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and surpasses the strongest single‑agent baseline by 24.0 points.
arXiv:2609.08402v1 Announce Type: cross
Abstract: Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned...
By Boao Yu, Zimo Chen, Junreng Rao, Yue Hu, Zhengqiu Zhu, Yong Zhao, Rusheng Ju
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
arXiv:2606. 20045v1 Announce Type: cross Abstract: UAV Vision-Language Navigation (UAV-VLN) is typically formulated as a holistic search-and-reach problem, where long-range target discovery and final target approach are optimized and evaluated jointly.
By Fanfu Xue, En Yu, Yantian Shen, Zhikun Hu, Hongjun Wang, Yang Yang, Xindi Wang, Jiande Sun
The paper introduces DroneCATS-Agent, a modular framework that places a multimodal large language model (MLLM) at the core of a drone’s control loop, allowing the model to decide actions solely from prompts. It presents the DroneCATS benchmark, evaluating MLLMs on four tasks—approaching, tracking, searching, and multi‑drone commanding—without fine‑tuning or function‑calling. Results show that while small open models can navigate reliably, they often fail by mismanaging protocol termination, highlighting a gap between perception and action planning in current MLLMs.
By Jaewoo Park, Minyoung Lee, Sukmin Seo, Moonbin Yim, Hyunwook Yoon, Dohoon Ryu, Daehee Kim, Myungseo Song, Jihyuk Byun, Seunggyu Chang, Taeho Kil, Jiseob Kim, Bado Lee, Geewook Kim
arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang