arXiv:2607. 24856v1 Announce Type: cross Abstract: Social media imagery (SMI) provides timely and fine-grained ground perspectives that are valuable for situational awareness and emergency response.
By Wenping Yin, Ziqi Liu, Naixia Mou, Weijia Li, Danfeng Hong, Hao Li
arXiv:2602.01163v2 Announce Type: replace
Abstract: Safe UAV emergency landing requires more than just identifying flat terrain; it demands understanding complex semantic risks (e.g., crowds, tempora...
By Chunliang Hua, Lei Zhang, Jiayang Sun, Chunlan Zeng, Xiao Hu
arXiv:2601.18493v2 Announce Type: replace
Abstract: Vision--language models (VLMs) show promise for disaster-response remote sensing, but existing benchmarks mainly emphasize scene-level or damage-ce...
By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.
By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
The paper introduces the Geo-Context Guided Visual Transformer, a model that augments remote sensing image analysis with geospatial embeddings and an asymmetric attention module. By converting heterogeneous geospatial variables into patch-aligned representations and assigning geospatial roles to attention heads, the approach improves disease prevalence prediction over existing vision-language and graph-based baselines. Ablation and visualization studies demonstrate its effectiveness and interpretability for health-related remote sensing tasks, especially when comprehensive geospatial data are scarce.
By Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
arXiv:2606. 18661v1 Announce Type: cross Abstract: Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios.
By Chengfu Liu, Dongyang Hou, Junwu Xiang, Cheng Yang, Xuezhi Cui, Zeyuan Wang, Liangtian Liu, Zelang Miao
DamageScope is a retrieval‑augmented framework that combines satellite imagery, Vision‑Language Models (VLMs), and Large Language Models (LLMs) to automate property damage assessment after natural disasters. It uses a Retrieval‑Augmented Generation (RAG) architecture to extract structured visual representations from satellite images, enabling interactive natural language queries. The system introduces a multi‑vector embedding‑based clustering algorithm that improves scalability and reduces indexing time by up to 14×, and a dual‑store data architecture that cuts LLM API calls, lowering operational cost and response latency by roughly 3×.
By Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas, Srimat T. Chakradhar
arXiv:2609.00046v1 Announce Type: cross
Abstract: Rapid and reliable disaster mapping of impacted areas, damaged infrastructure, and affected populations is essential for emergency response and recov...
By Yifan Yang, Lei Zou
arXiv:2608. 00012v1 Announce Type: cross Abstract: Multimodal Large Language Models (MLLMs) are increasingly used to interpret Earth observation data, yet their capability to support real-world disaster emergency response remains insufficiently evaluated.
By Fengxiang Wang, Qiuyang Yu, Yueying Li, Mingshuo Chen, Chengchi Fei, Kaiyi Xu, Lixin Gu, Wangxu Wei, Junchao Gong, Lipeng Ma, Jiong Wang, Fenghua Ling, Wenlong Zhang, Xue Yang, Wenjing Yang, Ben Fei, Long Lan
The paper introduces a two‑stage training framework that combines Supervised Fine‑Tuning (SFT) and Direct Preference Optimization (DPO) to improve multimodal disaster severity assessment. It creates two datasets—ReasoningSet for validated rationales and PreferenceSet for paired rationales—using a single Human‑in‑the‑Loop workflow. Experiments on InternVL‑3‑8B and LLaVA‑1.5‑7B show that SFT boosts classification accuracy and Macro‑F1, while DPO further enhances interpretability and alignment with human judgment.
By Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee, Fahad Khalid, Mourad Oussalah
arXiv:2601. 19099v2 Announce Type: replace-cross Abstract: Vision--language models (VLMs) achieve strong performance on many multimodal benchmarks but remain brittle on spatial reasoning tasks that require aligning abstract overhead representations with egocentric views.
By Yosub Shin, Michael Buriek, Igor Molybog