arXiv:2607. 24856v1 Announce Type: cross Abstract: Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response.
By Wenping Yin, Ziqi Liu, Naixia Mou, Weijia Li, Danfeng Hong, Hao Li
arXiv:2602.01163v2 Announce Type: replace
Abstract: Safe UAV emergency landing requires more than just identifying flat terrain; it demands understanding complex semantic risks (e.g., crowds, tempora...
By Chunliang Hua, Lei Zhang, Jiayang Sun, Chunlan Zeng, Xiao Hu
arXiv:2601.18493v2 Announce Type: replace
Abstract: Vision--language models (VLMs) show promise for disaster-response remote sensing, but existing benchmarks mainly emphasize scene-level or damage-ce...
By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.
By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
The paper introduces the Geo-Context Guided Visual Transformer, a model that augments remote sensing image analysis with geospatial embeddings and an asymmetric attention module. By converting heterogeneous geospatial variables into patch-aligned representations and assigning geospatial roles to attention heads, the approach improves disease prevalence prediction over existing vision-language and graph-based baselines. Ablation and visualization studies demonstrate its effectiveness and interpretability for health-related remote sensing tasks, especially when comprehensive geospatial data are scarce.
By Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui