arXiv AI

ABot-N1: Toward a General Visual Language Navigation Foundation Model

arXiv:2607. 10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks.

arXiv AI
Sep 1

LightNav-0: Eliciting VLM Spatial Intelligence for Generalist Embodied Navigation

arXiv:2608.30935v1 Announce Type: cross Abstract: Embodied navigation requires agents to translate heterogeneous goals and visual observations into actions across tasks, environments, and robot embod...

By Shaoan Wang, Aocheng Luo, Fei Huang, Jingyi Xu, Xiaoyang Wang, Yueyu Wang, Qianli Ma, Fan Yang, Ran Mei, Jia Wei, Jiangpeng Hu, Xuhao Liu, Hongming Chen, Yuanbin Shao, Yiyang Lin, Ziliang Li, Liang Pan, Xinhang Liu, Yuntao Ma, Tingxiang Fan
arXiv AI
Jun 9

SpaceVLN: A Zero-Shot Vision-and-Language Navigation Agent with Online Spatial Cognitive Memory and Reasoning

arXiv:2606. 08992v1 Announce Type: cross Abstract: Vision-and-Language Navigation in continuous environments requires agents to understand the spatial structure of previously unseen environments in order to follow language instructions.

By Yucheng Deng, Pingrui Lai, Xinhai Li, Chenjia Bai, Xiaoheng Deng, Chengnuo Sun, Xuelong Li, Hua Yang
arXiv Computer Vision
4d ago

InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.

By Hongpei Zheng, Hujun Yin
arXiv Computer Vision
Sep 14

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.

By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv Computer Vision
2d ago

Guide, Think, Act: Interactive Embodied Reasoning in Vision-Language-Action Models

The paper introduces GTA‑VLA, an interactive Vision‑Language‑Action framework that lets users guide robot policies with explicit visual cues such as affordance points, boxes, and traces. Unlike traditional direct sense‑to‑act models, GTA‑VLA incorporates a spatial‑visual Chain‑of‑Thought that blends human guidance with internal task planning, and couples this reasoning module with a lightweight reactive action head for efficient execution. Experiments on the SimplerEnv WidowX benchmark show a state‑of‑the‑art 81.2 % success rate, and the framework significantly improves task success under out‑of‑domain visual shifts and spatial ambiguities, demonstrating the benefit of interactive reasoning for failure recovery in embodied control.

By Yiran Ling, Qing Lian, Jinghang Li, Qing Jiang, Tianming Zhang, Xiaoke Jiang, Chuanxiu Liu, Jie Liu, Lei Zhang
arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv AI
Jul 3

CoFL-S: Spatially Queryable Sector Flow Fields for Local Language-Conditioned Navigation

arXiv:2607. 02222v1 Announce Type: cross Abstract: Vision-Language Navigation has increasingly emphasized high-level instruction reasoning, memory, global map construction, and instruction decomposition, while the low-level action representation remains comparatively underexplored.

By Haokun Liu, Zhaoqi Ma, Yicheng Chen, Wentao Zhang, Masaki Kitagawa, Zicen Xiong, Jinjie Li, Moju Zhao