GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.
By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin
arXiv:2608.21832v1 Announce Type: new
Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether mo...
By Md Abrar Jahin, Md Rizwan Parvez
arXiv:2606. 28397v1 Announce Type: cross Abstract: Vision-language navigation (VLN) has recently advanced with large language and multimodal models, enabling agents to follow natural-language instructions in unseen environments without training a task-specific navigation policy.
By Shaoxuan Li, Xiangyu Dong, Xiaoguang Ma, Junfeng Chen, Haoran Zhao, Yaoming Zhou
Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CIDQ), where a model must identify the nearest visible candidate to a unique reference object and estimate their gravity-aligned floor-plane distance.
arXiv:2606. 15427v1 Announce Type: cross Abstract: Spaceborne inspection systems often deploy perception models prior to launch, after which updating model weights or expanding fixed label sets becomes operationally impractical.
By Nicholas A. Welsh, Lennon J. Shikhman, Monty Nehru Attazs, Seemanthini K. Putane, Van Minh Nguyen, Ryan T. White
arXiv:2606. 17539v1 Announce Type: cross Abstract: Spatial VLMs have made substantial progress in geometric perception, yet complex spatial reasoning requiring multi-step inference over depth, distance, and scene relations remains challenging.
By Yatai Ji, An-Chieh Cheng, Yang Fu, Yukang Chen, Han Zhang, Zhaojing Yang, Wei Huang, Ka Chun Cheung, Song Han, Vidya Nariyambut Murali, Pavlo Molchanov, Jan Kautz, Simon See, Hongxu Yin, Ping Luo, Sifei Liu
EgoPathBench is a new dataset and benchmark that tests zero‑shot egocentric waypoint decision‑making in vision‑language models. Each task presents an egocentric RGB image, a natural‑language goal, and numbered visible waypoints, and models must return traversable candidates or an ordered route. The benchmark evaluates candidate feasibility, edge legality, and goal arrival under point‑agent or embodied geometry, covering 31,852 training, 1,345 validation, and 1,111 benchmark questions.
"whyItMatters":"The benchmark reveals that current VLMs perform poorly on integrated navigation tasks, highlighting a gap in spatial intelligence that can be addressed by fine‑tuning with the released training data."
By Yang Zhao, Zhuo Chen, Xubo Yang
In this paper, we tackle the Aerial Vision-and-Dialog Navigation (AVDN) task in the training-free setting for resource-efficient high-altitude UAV navigation. Naively applying MLLMs leads to unreliable navigation due to weak directional grounding and the lack of explicit spatial memory.
AnchorVLN is an open‑vocabulary vision‑language navigation system that separates semantic proposals from geometric metrics. It uses a VLM to generate semantics while a geometry module supplies reliable metric quantities such as range and bearing, all within a Model Context Protocol server. The system achieves 64.4% on instruction following and improves object‑reference accuracy, reducing median center error from 3.37 m to 2.48 m.
By Long Giang Vu, Chengkai Yao, Yuxin Liu, FNU Aryan, Rajath Chandrashekar Aralikatti
arXiv:2604.12102v3 Announce Type: replace
Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representati...
By Arun Sharma
Zero-shot waypoint navigation requires vision-language models to select, from the current first-person observation, a sequence of spatial actions that is feasible for the agent and reaches the goal, p...
RefineRank introduces a lightweight module, RefineNet, that jointly refines bounding boxes and ranks them for surgical spatio‑temporal grounding. By combining frozen medical vision‑language features with proposals from a frozen open‑set detector, it predicts coordinate corrections and quality scores for each candidate box, selecting the best refined or original box. On MedVidBench, RefineRank achieves the highest reported STG mIoU of 0.421 and improves oracle bounds and overall mIoU in controlled evaluations.
By Linzhe Jiang, Jiayuan Huang, Changhao Zhang, Chunyang Jiang, Zhehua Mao, Mobarak I. Hoque