arXiv:2607. 07292v1 Announce Type: cross Abstract: Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consistently across cities due to data-source heterogeneity and the lack of fine-grained semantic-temporal context in remote sensing data.
By Zeru Yang, Fang-Ying Gong, Steve H. L. Yim, Chau Yuen
arXiv:2606. 20167v1 Announce Type: new Abstract: Spatial prediction tasks are often limited by a lack of high-quality labelled ground-truth observations.
By Jonathan Hecht, Lukas Arzoumanidis, Ziyue Li, Youness Dehbi
Multimodal Large Language Models (MLLMs) have demonstrated substantial promise in spatial understanding. Existing works typically incorporate prior knowledge extracted from a pre-trained foundation model to further enhance the spatial awareness of MLLMs.
GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.
By Jinnao Li, Tingzhu Chen, Changbo Wang
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
arXiv:2605. 14925v2 Announce Type: replace-cross Abstract: Drone-view geo-localization aims to match a query drone image, often captured under adverse weather conditions (e.
By Yunsong Fang, Tingyu Wang, Zhedong Zheng
The paper introduces the Geo-Context Guided Visual Transformer, a model that augments remote sensing image analysis with geospatial embeddings and an asymmetric attention module. By converting heterogeneous geospatial variables into patch-aligned representations and assigning geospatial roles to attention heads, the approach improves disease prevalence prediction over existing vision-language and graph-based baselines. Ablation and visualization studies demonstrate its effectiveness and interpretability for health-related remote sensing tasks, especially when comprehensive geospatial data are scarce.
By Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu
The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.
By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv:2609.15305v1 Announce Type: cross
Abstract: Urban region representation learning commonly combines heterogeneous data sources, such as mobility flows, points of interest, and land-use informati...
By Sean Bin Yang, Ying Sun, Zongyi Xu, Tung Kieu, Jilin Hu, Bin Yang, Kristian Torp, Hua Lu, Torben Bach Pedersen
arXiv:2608. 03826v1 Announce Type: cross Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues.
By Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu
arXiv:2606.08918v2 Announce Type: replace
Abstract: Worldwide image geo-localization aims to determine where on Earth a single image was captured. However, visually similar scenes may lie thousands o...
By Junchao Cui, Xuanzi Ma, Wenqi Shi, Nan Wu, Biru Zhu, Xiangyang Luo