arXiv Computer Vision

NavHarness: Adaptive Goals for Agentic Vision-Language Navigation

arXiv AI
Sep 1

Scaffolding Foundation Models into Physical-World Agents Pushes the Frontier of Long-Horizon Navigation

The paper introduces NavMCP, a scaffolding framework that couples vision‑language models (VLMs) with navigation foundation models (NFMs) to enable long‑horizon physical‑world agents. NavMCP orchestrates three communication channels—intent, observation, and memory—to allow the VLM to decide what evidence to seek and the NFM to ground semantic sub‑goals into closed‑loop navigation, without retraining either model. The approach achieves state‑of‑the‑art results on several embodied question‑answering benchmarks and significantly outperforms episodic interfaces on the Unitree Go2 robot as task horizons lengthen.

By Zixing Lei, Gengze Zhou, Xiong-Hui Chen, Jiazhao Zhang, Yiyang Huang, Hang Yin, Haoqi Yuan, Qi Wu, Weixin Li, Siheng Chen
Hugging Face Trending Papers
Jul 2

Path-level Hindsight Instructions for Semantic Exploration in Vision-Language Navigation

On-policy exploration is a crucial component for training robust Vision-Language Navigation agents, as it exposes the policy to a broader state distribution. However, such exploration inevitably leads to trajectories that deviate from expert demonstrations, resulting in a semantic mismatch between the executed visual stream and the original language instruction.

arXiv Computer Vision
Sep 22

What do VLM-Based Vision-Language Navigation Models Rely on: Interpreting and Steering Policy Behavior

arXiv:2609.24576v1 Announce Type: cross Abstract: Modern Vision-Language Navigation (VLN) models rely mostly on pre-trained large Vision-Language Models (VLMs) to predict navigation actions. While th...

By D\'ebora Oliveira Makowski, Samiran Gode, Abhijeet Nayak, Marco Hutter, Cordelia Schmid, Lukas Rosenberger Schmid, Wolfram Burgard
arXiv AI
Sep 12

2AM: Grounding Agent-Side Memory as Guidance for Steerable Action Models in Long-Horizon Manipulation

The paper introduces 2AM, a system that separates memory and action execution in long‑horizon robot manipulation. 2AM stores task memory exclusively in a multimodal Agent, while a single RGB‑based, stateless Action Model performs motion based on language and optional 2D hints. On the LIBERO‑Mem benchmark, 2AM achieves 76.3% average completion—over 61 points higher than the best baseline—demonstrating that agent‑side memory and precise steering of the Action Model can substantially improve performance.

By Yutong Hu, Fengjiao Chen, Xuezhi Cao, Renaud Detry
arXiv AI
Sep 2

EmbodiedSkills: A Unified Framework for Orchestrating, Training, and Deploying VLA Agents

EmbodiedSkills is a unified framework that treats each skill decision as an execution proposal, checking prerequisites and verifying outcomes during long‑horizon vision‑language‑action tasks. It connects high‑level skill selection, bounded low‑level VLA execution, and post‑action verification through a fixed executable‑skill interface, enabling easy replacement of low‑level policies and recording of structured trajectories for supervision and adaptation. Instantiated with Qwen3‑VL and OpenPI/pi0.5 on RoboTwin 2.0 and LIBERO, the framework achieves high success rates (86.20% and 97.40% respectively) and demonstrates effective memory‑dependent task performance.

By Wei Wang, Wenqiao Zhang, Yutong Lin, Yuqian Yuan, Tianwei Lin, Jinhao Mao, Zhenxuan Fan, Mingjian Gao, Yang Dai, Wentong Li, Zheqi Lv, Zheng Dong, Yingjie Niu, Jiaqi Zhu, Jun Xiao, Chao Li, Yueting Zhuang
arXiv AI
2d ago

Seek Before You Move: Evidence Seeking for Progress Grounding in Vision-Language Navigation

SeekVLN is a new framework for Vision‑Language Navigation that addresses the problem of agents acting on insufficient evidence, termed Progress Myopia. It combines semantic progress reasoning with active evidence seeking, trained first with Future‑guided Reverse Generation to augment expert trajectories, and then refined via Counterfactual Contrastive Policy Optimization to reward beneficial seeking actions. Experiments on simulated benchmarks show significant gains, improving success rates by 12.7% on R2R‑CE and 7.5% on RxR‑CE, and real‑world tests demonstrate human‑like evidence‑seeking behavior.

By Zhimin Wang, Meiyuan Zhu, Duo Wu, Linjia Kang, Yajun Wang, Yuan Ni, Xiaohang Wang, Tianlu Pan, Jingyan Jiang, Yaowei Wang, Zhi Wang
arXiv AI
Jun 30

CLOSER-VLN: Closed-Loop Self-Verified Retrieval-Augmented Reasoning for Aerial Vision-Language Navigation

arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.

By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
arXiv AI
Jul 14

ABot-N1: Toward a General Visual Language Navigation Foundation Model

arXiv:2607. 10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks.

By Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
arXiv AI
Jun 8

Think Like a Pilot: Fine-Grained Long-Horizon UAV Navigation

arXiv:2606. 06836v1 Announce Type: cross Abstract: Language-guided UAV agents must execute long-horizon semantic instructions while producing smooth, physically feasible continuous flight commands, yet existing Vision-Language Navigation (VLN) benchmarks typically use discrete or coarse actions and existing UAV Vision-Language-Action (VLA) tasks focus on short, atomic maneuvers.

By Xiangyi Zheng, Xiangyu Wang, Qinan Liao, Zimu Tang, Yue Liao, Dongyue Lyu, Guodong Wang, Junjie Liu, Si Liu