arXiv:2608. 09101v1 Announce Type: cross Abstract: Semantic segmentation models are trained and evaluated against human-drawn masks, yet remote-sensing annotations are often coarse, incomplete, or misaligned; high overlap scores may then reflect agreement with imperfect labels rather than faithfulness to the image, creating an evaluation paradox.
By Shuaishuai Cao, Shuwei Peng, Meng Tang, Min Huang, Youjin Wang, Jie Chen, Jing Ouyang, Zhiwei Zhai
The paper introduces the Remote Sensing Copy-Move Question Answering (RSCMQA) task, aimed at interpreting complex tampering scenarios in land resource monitoring and national defense. It presents five datasets—RS-CMQA, RS-CMQA-B, Real-RSCM, RS-TQA, and RS-TQA-B—spanning 29 regions in 14 countries, and a region-discrimination-guided multimodal copy-move forgery perception framework (CMFPF) that improves question answering accuracy on tampered images. Experiments show CMFPF outperforms general VQA and RSVQA models, establishing a new benchmark for RSCMQA.
By Ze Zhang, Enyuan Zhao, Di Niu, Jie Nie, Xinyue Liang, Lei Huang
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
The paper introduces VertiCue-Bench, a diagnostic benchmark designed to test whether multimodal large language models (MLLMs) can perceive, ground, and utilize vertical structure information in remote-sensing natural scenes. It presents a three-stage framework—Perception, Grounding, Utilization—and a Representation Intervention Spectrum across various modalities to evaluate ten state-of-the-art models. The study finds a significant Vertical Structure Utilization Gap: while models show some geometric perception, they struggle to accurately link vertical evidence to spatial entities and incorporate it into semantic decisions.
By Jing Huang, Duanchu Wang, Junjie Yang, Zihang Cheng, Cheng Li, Lin Cui, Zhouyi Wu, Di Wang
Newly triggered landslides rarely carry immediate annotations, so cross-domain transferability determines the value of landslide mapping for emergency response and regional risk assessment. Vision foundation models have strengthened representational transfer, yet on unseen regions, events, and data sources they still generate high-confidence false alarms.
arXiv:2609.17269v1 Announce Type: new
Abstract: Multimodal large language models generate natural-language responses from visual inputs, yet may mention objects absent from an image. In medication as...
By Ziheng Ren, Qian Gao, Jun Fan, Guohui Ding, Zhenyu Yang, Yuteng Xiao
arXiv:2603. 04277v2 Announce Type: replace-cross Abstract: Autonomous aerial robots operating in GPS-denied or communication-degraded environments frequently lose access to camera metadata and telemetry, leaving onboard perception systems unable to recover the absolute metric scale of the scene.
By Yifei Chen, Chenqian Le, Jiayi Cheng, Xupeng Chen
GeoRefer-Bench is a new benchmark for verifiable geospatial referring segmentation that evaluates whether models correctly resolve spatial relations in overhead imagery. Each query is expressed as an executable logical form over a metric scene graph, and predictions are scored with Exact Query Success (EQS), requiring an exact match to the query’s referent set. The dataset contains 700 UAV scenes, 26,217 instances, 142,796 spatial relations, 20,916 executable queries across five reasoning levels, and additional paraphrases, unanswerable queries, counterfactual pairs, and leakage‑controlled splits.
By Shuaishuai Cao, Min Huang, Meng Tang, Xuan Liu, Youjin Wang, Hui Lin
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
By Jinnao Li, Tingzhu Chen, Changbo Wang
The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.
By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv:2604. 26051v2 Announce Type: replace-cross Abstract: The increasing number of satellites has improved the temporal resolution of Earth observation, making satellite-based flood mapping a promising approach for operational flood monitoring.
By Hyunho Lee, Wenwen Li
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo