The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.
By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv:2605.17949v2 Announce Type: replace
Abstract: Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the...
By Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Lang Sun, Jiaqi Liu, Jiashun Zhu, Jiasen Hu, Xu Na, Bo Yang
arXiv:2606. 02374v1 Announce Type: new Abstract: Earth Observation (EO) has fundamentally transformed the monitoring of environmental processes and human activities up to planetary scale.
By Steffen Knoblauch, Hao Li, Gengchen Mai, Konstantin Klemmer, Song Gao, WenWen Li
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.
arXiv:2606. 24997v1 Announce Type: new Abstract: Geographic implicit neural representations (INRs) learn to map any coordinate on Earth to a location embedding, implicitly encoding geospatial data into the weights of a neural network.
By Livia Betti, Sebastian Ricke, Ivica Obadic, Adam J. Stewart, Esther Rolf
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo
arXiv:2608. 03826v1 Announce Type: cross Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues.
By Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe perf...
arXiv:2512.02697v4 Announce Type: replace
Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-...
By Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, Liangpei Zhang
arXiv:2510. 13774v2 Announce Type: replace Abstract: Forecasting urban phenomena such as housing prices and public health indicators requires the effective integration of various geospatial data.
By Dominik J. M\"uhlematter, Lin Che, Ye Hong, Martin Raubal, Nina Wiedemann
Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes.
The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang