arXiv AI By Dongbin Na, Chanwoo Kim, Soonbin Rho, Giyun Choi, Gangbok Lee, Dooyoung Hong

Binary Tracking for Spatial QA and Navigation with Open Vision-Language Models

Read the original on arXiv AI →

arXiv:2606. 16902v1 Announce Type: cross Abstract: This work addresses spatial question answering for service robots traversing long egocentric routes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 16

EgoPathBench: Evaluating Zero-Shot Egocentric Waypoint Decision-Making in Vision-Language Models

EgoPathBench is a new dataset and benchmark that tests zero‑shot egocentric waypoint decision‑making in vision‑language models. Each task presents an egocentric RGB image, a natural‑language goal, and numbered visible waypoints, and models must return traversable candidates or an ordered route. The benchmark evaluates candidate feasibility, edge legality, and goal arrival under point‑agent or embodied geometry, covering 31,852 training, 1,345 validation, and 1,111 benchmark questions. "whyItMatters":"The benchmark reveals that current VLMs perform poorly on integrated navigation tasks, highlighting a gap in spatial intelligence that can be addressed by fine‑tuning with the released training data."

By Yang Zhao, Zhuo Chen, Xubo Yang
arXiv Computer Vision
Sep 14

AnchorVLN: Geometry-Anchored Vision-Language Grounding Reasoning for Open-Vocabulary Navigation

AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.

By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
Hugging Face Trending Papers
Jul 22

NavVerse: Benchmarking Indoor-to-Outdoor Embodied Navigation in Continuous Robot Simulation

Robots deployed in delivery, campus, and emergency-response settings often need to navigate from buildings to streets within a single continuous episode. Existing benchmarks usually evaluate indoor and outdoor navigation separately, and many abstract away robot execution, leaving exit finding, boundary traversal, adaptation, and kinodynamic failures underexplored.

arXiv AI
2d ago

UrbanVLA: A Vision-Language-Action Model for Urban Micromobility

UrbanVLA is a Vision‑Language‑Action framework designed to enable delivery robots to navigate large‑scale urban environments using long‑horizon route instructions. The model aligns noisy route waypoints with visual observations and plans trajectories, trained through a two‑stage pipeline of supervised fine‑tuning on simulated data and reinforcement fine‑tuning on mixed simulation and real‑world data. Experiments show UrbanVLA outperforms strong baselines by over 55% on the SocialNav task and demonstrates reliable real‑world navigation in large urban settings.

By Anqi Li, Zhiyong Wang, Jiazhao Zhang, Minghan Li, Yunpeng Qi, Zhibo Chen, Zhizheng Zhang, He Wang