arXiv AI

MapDream: Task-Driven Map Learning for Vision-Language Navigation

arXiv:2602. 00222v3 Announce Type: replace-cross Abstract: Vision-Language Navigation (VLN) requires agents to follow natural language instructions in partially observed 3D environments, motivating map representations that aggregate spatial context beyond local perception.

arXiv AI
Sep 17

Mem2Ego: Empowering Vision-Language Models with Global-to-Ego Memory for Long-Horizon Embodied Navigation

Mem2Ego introduces a vision‑language model for embodied navigation that combines global memory with egocentric visual inputs. By adaptively retrieving task‑relevant cues from a global memory module and aligning them with local perception, the framework improves spatial reasoning and decision‑making over long horizons. The method outperforms prior state‑of‑the‑art approaches on the HSSD and HM3D benchmarks and shows strong performance on a real robot.

By Lingfeng Zhang, Yuecheng Liu, Zhanguang Zhang, Matin Aghaei, Yixin Xiao, Yaochen Hu, Mohammad Ali Alomrani, David Gamaliel Arcos Bravo, Hongjian Gu, Zhiyuan Li, Yangzheng Wu, Zhanpeng Zhang, Raika Karimi, Atia Hamidizadeh, Guowei Huang, Haoping Xu, Tongtong Cao, Weichao Qiu, Xingyue Quan, Jianye Hao, Yuzheng Zhuang, Yingxue Zhang
arXiv AI
Jun 30

Pondering the Way: Spatial-perceiving World Action Model for Embodied Navigation

arXiv:2606. 29908v1 Announce Type: cross Abstract: Existing world model-based planners for visual navigation typically follow a verification-centric paradigm, decoupling goal intent from trajectory synthesis.

By Hong Chen, Daqi Liu, Zehan Zhang, Haiguang Wang, Tianhao Lu, Longfei Yan, Haiyang Sun, Fangzhen Li, Hongwei Xie, Bing Wang, Guang Chen, Hangjun Ye, Yihua Tan
arXiv Computer Vision
Sep 2

Qwen-Drive-1.0: An Initial Step towards a Vision-Language Foundation Model for Autonomous Driving

Qwen-Drive-1.0 is a vision‑language foundation model tailored for autonomous driving that builds on a pretrained VLM architecture. It incorporates a bird’s‑eye‑view perception head for 3D object detection, semantic occupancy prediction, and BEV map segmentation, and a Planning Expert that generates future ego trajectories from shared representations. Experiments show strong 3D perception, driving scene understanding, and competitive motion‑planning performance while largely preserving general vision‑language capabilities.

By Xin Zhou, Zongchuang Zhao, Zhibo Yang, Mingsheng Li, Humen Zhong, Shuai Bai, Du Chu, Ruizhe Chen, Zhaohai Li, Jun Tang, Qiuyue Wang, Mingkun Yang, Jiazhao Zhang, Dayiheng Liu, Dingkang Liang, Xiang Bai
arXiv AI
Aug 10

WNM-3D: A World Navigation Model with 3D Scene Conditioning for Closed-Loop VLN

arXiv:2608. 07267v1 Announce Type: new Abstract: Recent vision-language navigation (VLN) systems increasingly adapt pretrained vision-language models (VLMs) into vision-language-action (VLA) policies that map egocentric observations and language instructions directly to navigation actions.

By Yuehao Huang, Yunzi Wu, Xiaotao Zhang, Xinhai Li, Jiankun Dong, Jiajun Lv, Chi Zhang, Chenjia Bai, Yong Liu, Xuelong Li
arXiv Computer Vision
4d ago

InsightMap: Structured Spatial Modeling for Embodied Multimodal Reasoning

InsightMap is a framework that uses top‑down maps as explicit spatial memory and action‑conditioned prediction targets for language‑guided navigation. It links historical views to labeled map locations and employs a shared multimodal backbone to jointly learn navigation action prediction and post‑action map generation, providing auxiliary training supervision. The approach supports a unified RGB‑D pipeline for navigation, visual question answering, situated reasoning, and 3D grounding, achieving state‑of‑the‑art results on R2R‑CE, RxR‑CE, ScanQA, SQA3D, ScanRefer, and outperforming baselines on the Unitree Go2 platform.

By Hongpei Zheng, Hujun Yin
arXiv AI
Jul 14

ABot-N1: Toward a General Visual Language Navigation Foundation Model

arXiv:2607. 10383v1 Announce Type: cross Abstract: Visual Language Navigation foundation models aim to unify deep reasoning for grounded spatial decisions with broad versatility for diverse embodied tasks.

By Ruiyan Gong, Yingnan Guo, Junjun Hu, Jintao Kong, Xiaoxu Leng, Tianlun Li, Weize Li, Fei Liu, Zhicheng Liu, Jia Lu, Minghua Luo, Chenlin Ming, Yanfen Shen, Jiyue Tao, Zhengbo Wang, Mingyang Yin, Minqi Gu, Zihao Guan, Wei Guo, Guoqing Liu, Huachong Pang, Menglin Yang, Zeqian Ye, Xiaoxiao Geng, Zhining Gu, Honglin Han, Di Jing, Hongyu Pan, Mingchao Sun, Kuan Yang, Jianfang Zhang, Yanghong Chen, Ye He, Wei Mei, Jiahao Shi, Xiangpo Yang, Yanqing Zhu, Zedong Chu, Xiaolong Wu, Mu Xu
arXiv AI
Sep 21

Visual Navigation Transformer with Pose Attention

arXiv:2609.21212v1 Announce Type: cross Abstract: Learned navigation policies typically consume observations as a temporally ordered history, with positional encodings tying each observation to when...

By Beiming Li, Jaime Romero, Jonathan Diller, Vijay Kumar, Alejandro Ribeiro