EgoPathBench is a new dataset and benchmark that tests zero‑shot egocentric waypoint decision‑making in vision‑language models. Each task presents an egocentric RGB image, a natural‑language goal, and numbered visible waypoints, and models must return traversable candidates or an ordered route. The benchmark evaluates candidate feasibility, edge legality, and goal arrival under point‑agent or embodied geometry, covering 31,852 training, 1,345 validation, and 1,111 benchmark questions.
"whyItMatters":"The benchmark reveals that current VLMs perform poorly on integrated navigation tasks, highlighting a gap in spatial intelligence that can be addressed by fine‑tuning with the released training data."
By Yang Zhao, Zhuo Chen, Xubo Yang
AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.
By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, p...
Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.
arXiv:2606. 18634v1 Announce Type: cross Abstract: To locate a target object while exploring the unknown environment is a fundamental capability for autonomous agents, with applications ranging from search-and-rescue to field robots.
By Zecheng Yin, Benedict Jun Ma
UrbanVLA is a Vision‑Language‑Action framework designed to enable delivery robots to navigate large‑scale urban environments using long‑horizon route instructions. The model aligns noisy route waypoints with visual observations and plans trajectories, trained through a two‑stage pipeline of supervised fine‑tuning on simulated data and reinforcement fine‑tuning on mixed simulation and real‑world data. Experiments show UrbanVLA outperforms strong baselines by over 55% on the SocialNav task and demonstrates reliable real‑world navigation in large urban settings.
By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang
arXiv:2607. 20785v1 Announce Type: cross Abstract: Deploying navigation systems at scale requires a recipe that minimizes sensor assumptions, generalizes across robot embodiments, and trains efficiently.
By Arjun Majumdar, Avinash Sooriyarachchi, Benjamin Tibi, Chris Bamford, Elliot Chane-Sane, Guillaume Lample, Khyathi Raghavi Chandu, Ludovic Ho Fuh, Mathieu Poiree, Olivier Duchenne, Rosalie Millner, Srijan Mishra, Theo Cachet, Thomas Chabal
arXiv:2609.06476v1 Announce Type: cross
Abstract: Vision-and-Language Navigation in Continuous Environments (VLN-CE) requires an embodied agent to navigate unseen environments by following natural la...
By Shiqi Pan, Qi Zheng, Hanqin Sun, Youjian Zhang, Daquan Feng, Xu Wang
UrbanGround is a sandbox that tests how well multimodal large language model agents can translate local street‑view perception into reliable action within a physically realistic replica of Hong Kong. The platform offers closed‑loop first‑person interaction and an interactive map, allowing agents to navigate the 3D city and answer spatial questions. The study evaluates agents across three research questions—scene grounding, navigation over increasing distances, and robustness to route changes—revealing that while agents excel at visual recognition and short‑range reasoning, they struggle with sustained goal‑directed behavior and pedestrian‑aware movement.
By Tianjie Ju, Zheng Wu, Yueqing Sun, Yuhan Cui, Bobo Li, Shengqiong Wu, Pengzhou Cheng, Haodong Zhao, Zongru Wu, Xinbei Ma, Doris Zhang, Kunling Li, Mong-Li Lee, Wynne Hsu, Hao Fei, Qi Gu, Gongshen Liu, Zhuosheng Zhang
UniTrackPLA introduces a unified panorama-language-action model that simultaneously handles instruction‑guided navigation and dynamic person tracking for embodied robots. Its Panoramic‑Aware Encoding preserves azimuthal and temporal structure, allowing a shared vision‑language backbone to generate continuous waypoint chunks for both tasks. The model also employs World‑Action Consistency to predict future visual states and verify waypoint prefixes, enabling reliable action reuse and replanning when inconsistencies arise. A new OmniTrackNav‑Bench dataset and extensive real‑world experiments demonstrate significant performance gains over prior methods.
By Pengfei Qi, Haoran Lin, Sizhuang Chen, Kai Luo, Sirui Zhang, Xinqi Liu, Fei Cheng, Wenrui Chen, Liming Yin, Kailun Yang
The paper proposes a new architecture for large language model (LLM) agents that enhances spatial understanding by combining geometrical tools with an LLM orchestrator in grid‑world environments. It first gathers geodesic trajectories, vector‑quantizes them to create a representative subset, and then has the LLM label each trajectory with a natural language description, turning them into reusable tools. During operation, the LLM selects the appropriate tool based on the current state and goal, while low‑level control executes the chosen trajectory, enabling efficient decision‑making in a partially observable 2D grid setting.
By Gabriel Turinici
arXiv:2608.29483v1 Announce Type: cross
Abstract: Modern Vision-Language Models (VLMs) perform well above the human baseline in image geolocalization, a task critically important in disaster response...
By Arka Mukherjee, Soham Roy, Kartikeya Trivedi, Shreya Ghosh