Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CIDQ), where a model must identify the nearest visible candidate to a unique reference object and estimate their gravity-aligned floor-plane distance.
arXiv:2607. 03869v1 Announce Type: cross Abstract: Referring remote sensing image segmentation isolates the object named by a natural-language expression in an aerial image.
By Yuhang Jiang, Guohui Deng, Miaozhong Xu, Chao Ruan, Jinling Zhao, Linsheng Huang
arXiv:2604.12102v3 Announce Type: replace
Abstract: We describe compute-grounded reasoning (CGR), a design pattern in which code computes selected sub-problems from explicit intermediate representati...
By Arun Sharma
arXiv:2607. 00491v1 Announce Type: cross Abstract: Benchmarks for vision-language models (VLMs) mostly test observational spatial reasoning: models describe relations already visible in the input.
By Leyuan Yu, Xiao Tang, Minghao Liu, Xinyuan Li, Xiaokai Bai, Sheng Zhou, Qunshu Lin, Weihao Xuan, Naoto Yokoya
arXiv:2608.21832v1 Announce Type: new
Abstract: Computer-use agents ground natural-language instructions in screenshots to locate interface elements, yet existing benchmarks do not isolate whether mo...
By Md Abrar Jahin, Md Rizwan Parvez
GeoContext is a new vision‑language geolocation benchmark that introduces two tasks: GeoHint, where a model must localize an image given a coarse location hint, and GeoVerify, where a model must decide if an image was taken within 150 m of a claimed place. The benchmark builds a context ladder by stratifying nearby reference points by distance and referenceability, allowing the same image to be evaluated under varying context. Evaluation of five models on 109 sites in 30 cities shows that hint repetition is low, localization error grows with hint distance, and models struggle to achieve high discriminability in GeoVerify, with many false acceptances reported with high confidence.
By Yifan Zhang, Kai Wang
arXiv:2601. 19099v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views.
By Yosub Shin, Michael Buriek, Igor Molybog
ARGOS is a new benchmark and agent framework for multi‑camera person search that transforms the task from one‑shot retrieval into interactive reasoning using partial witness clues. The framework requires agents to plan questions, use spatial or temporal tools, and interpret ambiguous natural‑language responses within a limited turn budget, leveraging a Spatio‑Temporal Topology Graph that encodes camera connectivity and transition times. The benchmark includes 2,691 tasks across 14 real‑world scenarios, divided into semantic, spatial, and temporal tracks, and introduces Turn‑Weighted Success (TWS) as a metric that jointly measures correctness and turn efficiency, with current best agents achieving TWS scores of 0.383 and 0.590 on the spatial and temporal tracks respectively.
By Myungchul Kim, Kwanyong Park, Junmo Kim, In So Kweon
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo
arXiv:2606. 27876v1 Announce Type: cross Abstract: Spatial intelligence is essential for low-altitude unmanned aerial vehicle (UAV) perception, collaboration, and navigation.
By Haoyu Zhang, Meng Liu, Qianlong Xiang, Kun Wang, Yaowei Wang, Liqiang Nie
arXiv:2605.17630v3 Announce Type: replace
Abstract: Frozen segmentation foundation models often fail when the target class appears in a form that is weakly represented during pretraining. To address...
By Abderrahmene Boudiaf, Irfan Hussain, Sajid Javed