arXiv AI By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin

GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

Read the original on arXiv AI →

GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Aug 3

SpatialQuery: Benchmarking Geometry-Grounded Multi-Instance Spatial Reasoning in Vision-Language Models

Vision-language models (VLMs) achieve strong semantic understanding but remain unreliable in metric spatial reasoning, particularly when queries require comparing multiple instances of the same object category. We study this problem through the Closest-Instance Distance Query (CIDQ), where a model must identify the nearest visible candidate to a unique reference object and estimate their gravity-aligned floor-plane distance.

arXiv AI
Sep 10

GeoContext: One Context Ladder, Two Failure Modes in Vision-Language Geolocation: Flat Reliance on User-Provided Location Context and False Confirmation of Location Claims

GeoContext is a new vision‑language geolocation benchmark that introduces two tasks: GeoHint, where a model must localize an image given a coarse location hint, and GeoVerify, where a model must decide if an image was taken within 150 m of a claimed place. The benchmark builds a context ladder by stratifying nearby reference points by distance and referenceability, allowing the same image to be evaluated under varying context. Evaluation of five models on 109 sites in 30 cities shows that hint repetition is low, localization error grows with hint distance, and models struggle to achieve high discriminability in GeoVerify, with many false acceptances reported with high confidence.

By Yifan Zhang, Kai Wang