arXiv Machine Learning

GeoArbiter: Verifiability-Guided Grounding for Remote-Sensing Multimodal LLMs

arXiv:2608. 00877v1 Announce Type: new Abstract: Remote-sensing multimodal large language models (MLLMs) often assert facts that imagery cannot establish, such as a facility's identity or function.

arXiv Machine Learning
Aug 11

Contrastive Mask Fidelity: Reference-Free Auditing of Ground-Truth Masks in Remote Sensing Semantic Segmentation

arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.

By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
arXiv Computer Vision
Sep 24

Copy-Move Forgery Detection and Question Answering for Remote Sensing Image

The paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task, aimed at interpreting complex tampering scenarios in land resource monitoring and national defense. It presents five datasets—RS-CMQA, RS-CMQA-B, Real-RSCM, RS-TQA, and RS-TQA-B—spanning 29 regions in 14 countries, and a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF) that improves question answering accuracy on tampered images. Experiments show CMFPF outperforms general VQA and RSVQA models, establishing a new benchmark for RSCMQA.

By Ze Zhang, Enyuan Zhao, Di Niu, Jie Nie, Xinyue Liang, Lei Huang
arXiv Computer Vision
Sep 24

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.

By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv Computer Vision
Sep 3

Beyond Appearance: Can Multimodal Large Language Models Exploit Vertical Structure for Remote Sensing Natural Scene Understanding?

The paper introduces VertiCue-Bench, a diagnostic benchmark designed to test whether multimodal large language models (MLLMs) can perceive, ground, and utilize vertical structure information in remote-sensing natural scenes. It presents a three-stage framework—Perception, Grounding, Utilization—and a Representation Intervention Spectrum across various modalities to evaluate ten state-of-the-art models. The study finds a significant Vertical Structure Utilization Gap: while models show some geometric perception, they struggle to accurately link vertical evidence to spatial entities and incorporate it into semantic decisions.

By Jing Huang, Duanchu Wang, Junjie Yang, Zihang Cheng, Cheng Li, Lin Cui, Zhouyi Wu, Di Wang
Hugging Face Trending Papers
Aug 10

GeoPhysAdapter: Scale-Matched Geophysical Adaptation for Cross-Domain Landslide Mapping with Vision Foundation Models

Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still generate high-confidence false alarms.

arXiv AI
Jul 14

VehAnchor: Metadata-Free Metric Scale Recovery from Vehicle Cues in Aerial Imagery

arXiv:2603. 04277v2 Announce Type: replace-cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene.

By Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
arXiv AI
Sep 25

GeoRefer-Bench: A Benchmark from Referring Pixels to Verifiable Geospatial Reasoning

GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.

By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin
arXiv Computer Vision
Aug 27

GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction

GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.

By Jinnao Li, Tingzhu Chen, Changbo Wang
arXiv AI
2d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li