arXiv AI

Post-Disaster Affected Area Segmentation with a Vision Transformer (ViT)-based EVAP Model using Sentinel-2 and Formosat-5 Imagery

arXiv:2507. 16849v3 Announce Type: replace-cross Abstract: We propose a vision transformer (ViT)-based deep learning framework to refine disaster-affected area segmentation from remote sensing imagery, aiming to support and enhance the Emergent Value Added Product (EVAP) developed by the Taiwan Space Agency (TASA).

arXiv AI
2d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv Computer Vision
Sep 21

DisasterInsight: A Building-Centric Benchmark for Evaluating Vision--Language Models in Disaster Response

DisasterInsight is a building‑centric benchmark designed to evaluate vision‑language models (VLMs) for disaster response. Built on the xBD satellite dataset, it adds OpenStreetMap‑derived functional labels to 134,108 building instances and offers 15 task types, including instance assessment, scene counting, multi‑instance reasoning, and structured report generation. Experiments show that VLMs excel at visible damage detection but struggle with building function, multi‑instance reasoning, counting, and grounded reporting, and instruction tuning only partially mitigates these gaps.

By Sara Tehrani, Yonghao Xu, Leif Haglund, Amanda Berg, Gulnaz Zhambulova, Michael Felsberg
arXiv Computer Vision
Sep 3

Vision-Language Model for Accurate Crater Detection

The paper presents a deep‑learning crater detection algorithm (CDA) based on the OWLv2 Vision Transformer, fine‑tuned with Low‑Rank Adaptation on a manually labeled IMPACT dataset. It optimizes a combined loss of CIoU for localization and contrastive loss for classification, achieving a maximum recall of 92.6% and precision of 71.4% on lunar images. The method demonstrates reliable crater detection under varied illumination and rugged terrain, supporting safer lunar landings.

By Patrick Bauer, Marius Schwinning, Florian Renk, Andreas Weinmann, Hichem Snoussi
arXiv AI
Jun 26

On-board Remote-Sensing Foundation Models for Unsupervised Change Detection of Disaster Events

arXiv:2606. 27018v1 Announce Type: cross Abstract: Remote Sensing Foundation Models (RSFMs) have emerged as a powerful alternative to supervised models for Earth Observation, allowing satellites to autonomously trigger high-resolution captures or adjust tasking parameters upon detecting an anomaly, thereby maximizing the utility of the mission's limited power and computational resources.

By S. Ram\'irez-Gallego
arXiv AI
Jun 2

CAFOSat: A Strongly Annotated Dataset for Infrastructure-Aware CAFO Mapping Using High-Resolution Imagery

arXiv:2606. 00548v1 Announce Type: cross Abstract: Concentrated Animal Feeding Operations (CAFOs) play an important role in agricultural production but are also associated with environmental, public health, and disease surveillance concerns.

By Oishee Bintey Hoque, Nibir Chandra Mandal, Mandy L Wilson, Samarth Swarup, Madhav Marathe, Abhijin Adiga
arXiv Computer Vision
Aug 28

Mapping Woody Vegetation from Multi-Source Imagery and Prediction Fusion for Enhanced Data Efficiency and Accuracy

The paper presents a framework that enhances deep‑learning tree‑cover mapping in New South Wales by fusing multiple imagery sources and normalizing image quality. It introduces an image‑composition technique that removes defects and a prediction‑fusion method that reduces reliance on any single image, together cutting errors by 38.2 % and 53.6 % respectively. Label transfer across diverse imagery further boosts data efficiency, yielding error reductions of 28.1 %–76.2 % and a 13‑fold decrease in performance variability across dates.

By Kal Backman, Jared Wood, Adam Roff