arXiv:2608. 12308v1 Announce Type: cross Abstract: Aerial vision-language navigation (VLN) requires an embodied agent to integrate visual evidence over time, plan future actions, and determine when it has reached a navigation goal under partial observability.
By Yan Deng, Fei Xu
arXiv:2608. 07557v1 Announce Type: cross Abstract: Vision-Language Navigation for Unmanned Aerial Vehicles (UAV-VLN) requires rapid and reactive control in complex 3D environments.
By Peng Xu, Chengcheng Wang, Shaohua Wan
arXiv:2606. 14772v1 Announce Type: cross Abstract: Aerial Embodied Question Answering (EQA) requires Unmanned Aerial Vehicles (UAVs) to actively perceive the environment and answer natural language questions.
By Wenhao Lu, Zhengqiu Zhu, Xiaofeng Wang, Xiaoran Zhang, Yatai Ji, Yong Zhao, Yue Hu, Yingzhen Nie, Jinlong Zhu, Zheng Zhu
Air-Ground Collaborative Vision-and-Language Navigation (AGC-VLN) pairs an unmanned aerial vehicle (UAV) with a global bird’s‑eye view and an unmanned ground vehicle (UGV) with a local first‑person view, creating a shared bird’s‑eye map that enables collaboration. The training‑free baseline decomposes navigation into VLM‑based semantic reasoning and deterministic geometric execution, allowing the UAV to render the UGV’s pose and target markers while the UGV plans a road‑following path using the shared map. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and surpasses the strongest single‑agent baseline by 24.0 points.
Air-Ground Collaborative Vision-and-Language Navigation (AGC‑VLN) pairs a UAV with a global bird’s‑eye view and a UGV with a local first‑person view, creating a shared bird’s‑eye map that displays the UGV’s pose and the target location. The UAV localizes the target in its downward view and flies toward it, while the UGV uses the shared map to plan a road‑following path with a frozen VLM and execute it under closed‑loop control. In CARLA‑Air’s Town10HD scene, AGC‑VLN achieves a 77.0% joint success rate, a 27.0% improvement over the weaker individual agent and outperforms the strongest single‑agent baseline by 24.0 points.
By Shuning Zhang, Liang Li, Yunheng Wang, Tao Wang, Yihang Kang, Renjing Xu
arXiv:2606. 12402v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) are increasingly deployed as high-level planners for embodied agents, with an emerging strategy of scaling test-time compute to improve capability.
By Jadelynn Dao, Milan Ganai, Yasmina Abukhadra, Ajay Sridhar, Mozhgan Nasr Azadani, Katie Luo, Clark Barrett, Jiajun Wu, Chelsea Finn, Marco Pavone
arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang
arXiv:2606. 06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers.
By Xiangyi Zheng, Xiangyu Wang, Qinan Liao, Zimu Tang, Yue Liao, Dongyue Lyu, Guodong Wang, Junjie Liu, Si Liu
arXiv:2607. 21400v1 Announce Type: cross Abstract: Vision-and-Language Navigation (VLN) enables embodied agents to follow natural-language instructions.
By Jiabin Lou, Haopeng Wang, Yuanshuai Wang, Xinyu Liu, Xuxin Lv, Yuxin Guo, Lei Huang, Rongye Shi, Wenjun Wu
arXiv:2608. 11739v1 Announce Type: cross Abstract: The prevailing recipe for Vision-Language-Action (VLA) models couples a pretrained VLM with a separately trained flow-matching action expert.
By Yicheng Liu, Zibin Dong, Baijun Ye, Tianyuan Yuan, Tao Jiang, Anqi Yang, Shicheng Cao, Haonan Liu, Yue Sun, Zihan Guo, Xiao Liu, Dong Ke, Changxun Pan, Chenru Wu, Tailai Cheng, Xiaoshu Ren, Xinlei Zhang, Jianning Cui, Zijie Zhao, Haoyu Zhang, Kaiming Xu, Haodong Yang, Bowen Zhang, Jiahui Niu, Shaoting Zhu, Shiduo Zhang, Hang Zhao
arXiv:2609.08402v1 Announce Type: cross
Abstract: Air-Ground Object Search (AGOS) in urban environments is a challenging embodied task, which requires an Unmanned Aerial Vehicle (UAV) and an Unmanned...
By Boao Yu, Zimo Chen, Junreng Rao, Yue Hu, Zhengqiu Zhu, Yong Zhao, Rusheng Ju
arXiv:2607. 08359v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables UAV autonomous navigation in unknown environments by mapping language instructions to real-time visual inputs.
By Xueke Zhu, Qingyan Meng, Liutao Yu, Wei Zhang, Zhengyu Ma, Huihui Zhou, Yonghong Tian