arXiv Computer Vision

A Multimodal Large Language Model-Driven Framework for Context-Aware UAV Emergency Landing Site Selection

arXiv AI
1d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv AI
Jul 20

Human-Inspired Neuro-Symbolic World Modeling and Logic Reasoning for Interpretable Safe UAV Landing Site Assessment

arXiv:2510. 22204v3 Announce Type: replace-cross Abstract: Reliable assessment of safe landing sites in unstructured environments is essential for deploying Unmanned Aerial Vehicles (UAVs) in real-world applications such as delivery, inspection, and surveillance.

By Weixian Qian, Tianyi Yang, Sebastian Schroder, Yao Deng, Jiaohong Yao, Xiao Cheng, Richard Han, Xi Zheng
arXiv AI
Jun 18

LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

arXiv:2606. 18661v1 Announce Type: cross Abstract: Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios.

By Chengfu Liu, Dongyang Hou, Junwu Xiang, Cheng Yang, Xuezhi Cui, Zeyuan Wang, Liangtian Liu, Zelang Miao
arXiv AI
3d ago

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.

By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv AI
Sep 10

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots. The method also employs object‑conditioned visual reasoning and egocentric regional guidance, achieving state‑of‑the‑art success rates on the UAV‑ON benchmark.

By Jianqiang Xiao, Xiang Deng, Yuexuan Sun, Yanjin Wu, Wenbiao Yan, Liqiang Nie
Hugging Face Trending Papers
Sep 8

Dual-Layer Semantic-Spatial Belief Mapping for Aerial Object Goal Navigation

The paper introduces AeroBelief, a dual‑layer semantic‑spatial belief mapping framework for aerial object goal navigation. It separates broad contextual plausibility (intuition layer) from target‑specific evidence (evidence layer) and fuses them into persistent spatial belief hotspots, while also employing object‑conditioned visual reasoning and temporally stable regional guidance. Experiments on the UAV‑ON benchmark show AeroBelief outperforms prior methods in success rate, object success rate, and SPL.

Hugging Face Trending Papers
Jun 17

LandslideAgent with Multimodal LandslideBench: A Domain-Rule-Augmented Agent for Autonomous Landslide Identification and Analysis

Intelligent landslide hazard interpretation is critical for disaster prevention, yet current paradigms struggle to simultaneously extract visual features and high-level geoscientific semantics, while general-purpose vision-language models (VLMs) suffer from perceptual limitations and domain hallucinations in complex geological scenarios. To address these challenges, we propose an instruction-driven agentic framework comprising three components.

arXiv AI
Jun 6

DisasterBench: A Multimodal Benchmark for UAV-Based Disaster Response in Complex Environments

arXiv:2606. 06217v1 Announce Type: cross Abstract: When a disaster unfolds, responders must answer not only what is happening, but also why it is happening, what will happen next, and what to do now, often from noisy low-altitude UAV views and under tight on-site compute constraints.

By Tan Zhang, Quanyou Li, Lu Zhang, Jun Liu, Xiaofeng Zhu, Ping Hu
arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
arXiv AI
Sep 10

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.

By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer