Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, p...
EgoPathBench is a new dataset and benchmark that tests zero‑shot egocentric waypoint decision‑making in vision‑language models. Each task presents an egocentric RGB image, a natural‑language goal, and numbered visible waypoints, and models must return traversable candidates or an ordered route. The benchmark evaluates candidate feasibility, edge legality, and goal arrival under point‑agent or embodied geometry, covering 31,852 training, 1,345 validation, and 1,111 benchmark questions.
"whyItMatters":"The benchmark reveals that current VLMs perform poorly on integrated navigation tasks, highlighting a gap in spatial intelligence that can be addressed by fine‑tuning with the released training data."
By Yang Zhao, Zhuo Chen, Xubo Yang
arXiv:2609.27076v1 Announce Type: new
Abstract: Open-vocabulary visual grounding enables robots to localise task-relevant entities from natural-language queries without dependence on predefined perce...
By Linus Nwankwo, Muslim Alaran, Christian Rauch, Stanley Chukwuebuka Obilikpa, Elmar Rueckert
arXiv:2606. 30632v1 Announce Type: cross Abstract: Can the robot use a plate to cut a cake if no knife is available?
By Yuhong Deng, Yuyao Liu, David Hsu
AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.
By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv:2607. 18200v1 Announce Type: cross Abstract: Robots in cluttered indoor spaces often fail not because they cannot generate collision-free paths, but because a fixed safety margin is mis-calibrated: conservative margins cause detours and timeouts, while permissive margins lead to near-boundary shortcuts under perception bias.
By Junyi Hu, Shuaihang Yuan, Geeta Chandra Raju Bethala, Anthony Tzes, Yi Fang
As Multimodal Large Language Models (MLLMs) are increasingly deployed in decision-critical pipelines such as robotics, embodied AI, and safety monitoring, the opacity of their spatial judgments limits operator trust and auditability. MLLMs demonstrate strong reasoning but often struggle with fine-grained spatial understanding and object hallucination.
arXiv:2605. 20306v2 Announce Type: replace-cross Abstract: We introduce WildRoadBench, a wild aerial road-damage grounding benchmark that couples direct visual grounding by vision-language models with autonomous research-and-engineering by LLM-driven agents on a single professionally annotated UAV corpus.
By Bingnan Liu, Chenhang Cui, Rui Huang, Jiani Luo, Zhirong Shen, Tinghao Wang, Xiande Huang, Lingbei Meng, Fei Shen, An Zhang
Vision-language-action (VLA) models enable robot navigation from natural language and visual goals, but remain susceptible to perceptual distractions and ambiguous scene interpretations. This paper presents the first empirical evaluation of visual grounding for VLA navigation policies.
arXiv:2606. 00095v1 Announce Type: cross Abstract: Vision-Language Navigation (VLN) enables embodied agents to reach target locations in unseen environments by following language instructions.
By Kailing Li, Tianwen Qian, Lijin Yang, Yuqian Fu, Jingyu Gong, Xiaoling Wang, Liang He
arXiv:2607. 06882v1 Announce Type: cross Abstract: Visual navigation policies built on large pretrained models have so far followed a common recipe: a dedicated visual encoder, a bespoke action head, and training on thousands of hours of cross-embodiment datasets.
By Peter Bohm, Saimunur Rahman, Abdelwahed Khamis, Sagun Man Singh Shrestha, Chris McCool, Peyman Moghadam
PointRL introduces a verifiable reinforcement learning framework that learns point-level vision‑language grounding from heterogeneous annotation evidence such as bounding boxes, masks, and instance labels. The method converts these annotations into pointing instructions while preserving target supports, instance membership, and set constraints as hidden verifier evidence, which a deterministic checker uses to score predictions. Evaluation on PointArena shows that PointRL improves Qwen3.5‑4B’s accuracy from 56.11% to 65.58%, and similar gains are observed on RoboSpatial, BLINK, and Ref‑Adv benchmarks.
By Jingyang Su, Pu Cao, Xiuze Jin, Longyue Zhang, Qing Song, Lu Yang