SatNav is a new, scalable benchmark for long‑horizon vision‑language navigation (VLN) with unmanned aerial vehicles (UAVs), built from high‑resolution satellite imagery. It generates 118,000 navigation episodes across 59 scenes in 18 cities, using satellite crops to approximate UAV nadir views and featuring three task families—Boundary, Landmark, and Route—to test long‑term memory and geospatial reasoning. The benchmark also introduces SwiftVLN, a modular framework for memory component experimentation, and demonstrates that models trained on satellite data can transfer to real‑flight UAV observations.
By Jiajun Jiang, Chunliang Hua, Zichun Chen, Yanxing Wu, Zeyuan Yang, Jie Song, Xiao Hu
arXiv:2609.00920v1 Announce Type: cross
Abstract: Vision-and-Language Navigation (VLN) requires an agent to navigate through unseen 3D environments according to natural-language instructions. Explici...
By Zhixin Wang, Chengzheyi Yao, Leyuan Liu, Xiaosong Zhang, Yongzhao Zhang
arXiv:2606. 08992v1 Announce Type: cross Abstract: Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions.
By Yucheng Deng, Pingrui Lai, Xinhai Li, Chenjia Bai, Xiaoheng Deng, Chengnuo Sun, Xuelong Li, Hua Yang
V-Link is a method designed to enhance Vision‑Language‑Action (VLA) models by recovering visual representations during the transfer from vision‑language (VL) features to action (A) features. It introduces complementary Spatial and Semantic Query representations that are injected into Action DiT through asymmetric pathways, providing both semantic augmentation and dedicated geometric conditioning for action generation. Experiments on LIBERO, LIBERO‑Plus, RoboTwin 2.0, and real‑world AGIBOT A3 Ultra tasks show significant performance gains over the base GR00T N1.6 model.
By Yehao Lu, Jiarui Yang, Yuning Su, Yufeng Xie, Yu Zhong, Yazhou Zhang, Haiyu Lan, Kaixiang Lu, Peiwen Lin, Chuang Wang, Zequn Qin, Enyu Li, Xi Li
arXiv:2603.26788v3 Announce Type: replace-cross
Abstract: Zero-shot object navigation requires agents to locate unseen targets in unfamiliar environments without prior maps or task-specific training....
By Feng Wu, Wei Zuo, Wenliang Yang, Jun Xiao, Yang Liu, Xinhua Zeng
Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first empirical evaluation of visual grounding for VLA navigation policies.