arXiv AI

MCANet: A Multi-Scale Class-Specific Attention Network for Multi-Label Post-Hurricane Damage Assessment Using UAV Imagery

arXiv AI
Jul 28

Advancing All-Weather Building Damage Mapping to the Instance Level: Outcomes and Insights from the 2026 Bright Challenge

arXiv:2607. 22746v1 Announce Type: cross Abstract: Rapid post-disaster response requires timely, building-level information on whether structures remain intact, are damaged, or are destroyed.

By Hongruixuan Chen, He Huang, Haifeng Wang, Jian Song, Junjue Wang, Weihao Xuan, Hamish Mitchell, Jiepan Li, Wei He, Liangpei Zhang, Zijie Wang, Chen Zhong, Jiazhen Zhao, Lei Hu, Ting Hu, Hongyan Zhang, Gregory Angelides, Miriam Cha, Clifford Broni-Bediako, Junshi Xia, Taylor Perron, Naoto Yokoya
arXiv AI
Jun 2

Attention mechanisms and transfer learning for robust peach leaf damage classification under domain shift

arXiv:2606. 02045v1 Announce Type: cross Abstract: Artificial intelligence provides a practical framework for crop damage assessment from imagery data, supporting early decision-making in agricultural management.

By Adri\'an C\'anovas-Rodriguez, Miguel A. Gonz\'alez-Ill\'an, Maria Fernanda Garc\'ia-Cruz, Pedro Nortes Tortosa, Jos\'e Salvador Rubio-Asensio, Miguel A. Zamora Izquierdo, Juan Antonio Mart\'inez Navarro, Antonio F. Skarmeta
arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
arXiv AI
Sep 2

Towards reliable multimodal disaster severity assessment through preference optimization and explainable vision-language reasoning

The paper introduces a two‑stage training framework that combines Supervised Fine‑Tuning (SFT) and Direct Preference Optimization (DPO) to improve multimodal disaster severity assessment. It creates two datasets—ReasoningSet for validated rationales and PreferenceSet for paired rationales—using a single Human‑in‑the‑Loop workflow. Experiments on InternVL‑3‑8B and LLaVA‑1.5‑7B show that SFT boosts classification accuracy and Macro‑F1, while DPO further enhances interpretability and alignment with human judgment.

By Yuanjun Zhang, Fuzel Ahamed Shaik, Suvojit Acharjee, Fahad Khalid, Mourad Oussalah
arXiv AI
Jul 16

Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery

arXiv:2507. 16849v3 Announce Type: replace-cross Abstract: We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery, aiming to support and enhance the Emergent Value Added Product (EVAP) developed by the Taiwan Space Agency (TASA).

By Yi-Shan Chu, Hsuan-Cheng Wei
arXiv Computer Vision
Sep 11

Vision Transformer-Based Multi-Level Feature Fusion for Multi-Label Sewer Defect Classification

The paper introduces Sewer-Transformer-ML, a hierarchical vision Transformer that fuses multi‑level features for multi‑label sewer defect classification, and two lightweight variants, Sewer-MobileNet-ML and Sewer-Mobile-TransNet, tailored for resource‑constrained inspection scenarios. On the Sewer‑ML test set, Sewer‑Transformer‑ML‑Base achieved an $F2_{ ext{CIW}}$ of 65.68% and an $F1_{ ext{Normal}}$ of 92.68%, topping the public leaderboard and surpassing the next best method by 7.6 percentage points in $F2_{ ext{CIW}}$. The lightweight Sewer‑MobileNet‑ML reached a comparable $F2_{ ext{CIW}}$ of 65.73% with only 17 M parameters, a 95% reduction from the base model, while Sewer‑Mobile‑TransNet achieved 96.43% accuracy under the standard data split, and ablation studies highlighted the effectiveness of direct concatenation for Transformer features and attention‑based fusion for multiscale CNN features.

By Xu Fang, Zhuoran Wang, Qing Li, Shengyu Zhang, Guanzhi Deng, Jianbiao He, Qingquan Li