arXiv Computer Vision

A Large Scale Open-Source Image and Video Dataset for Robust Wildfire Detection and Classification

arXiv AI
Sep 1

Toward Generalizable Deep Learning Based Peatland Fire Detection via Walsh Hadamard Transform and Domain Adaptation

arXiv:2603.02465v2 Announce Type: replace-cross Abstract: Machine learning-based wildfire detection has advanced significantly using deep learning models trained on large wildfire image and video dat...

By Emadeldeen Hamdan, Ahmad Faiz Tharima, Mohd Zahirasri Mohd Tohir, Dayang Nur Sakinah Musa, Erdem Koyuncu, Adam J. Watts, Ahmet Enis Cetin
arXiv Computer Vision
Sep 2

Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery

The paper proposes a data-driven protocol that uses multispectral Landsat‑8 imagery and connected‑component analysis to characterize fire‑region size distributions for active wildfire segmentation. Three segmentation architectures—U‑Net, DeepLabV3+, and SegFormer—are evaluated under different SWIR‑based spectral configurations, with U‑Net showing the strongest robustness and SWIR2 consistently delivering the best results. The study highlights the importance of both spectral band selection and architectural design for robust satellite‑based active wildfire mapping, especially when training on low fire‑pixel density images.

By Matheus F. Kovaleski, Cristiano Premebida, Jo\~ao Ruivo Paulo
arXiv Computer Vision
Sep 2

Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3

The paper examines RGB‑infrared fusion for binary wildfire segmentation using UAV imagery on the FLAME3 dataset. It compares RGB and infrared baselines with three fusion strategies across U‑Net, DeepLabV3+, and SegFormer architectures. Results show thermal data dominates segmentation performance, and feature‑level fusion with transformer‑based models yields the best results.

By Matheus F. Kovaleski, Lu\'is Garrote, Cristiano Premebida, J\'er\^ome Mendes, Jo\~ao Ruivo Paulo
arXiv AI
Sep 10

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.

By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer
Hugging Face Trending Papers
Jul 8

ASFR-Net: Adversarial Alignment and Spatio-Frequency Refinement Network for Heterogeneous Remote Sensing Image Change Detection

The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.

arXiv Machine Learning
2d ago

A Sentinel-2 benchmark dataset for deep-learning active-fire segmentation across 25 California wildfires

The article introduces an open image dataset for active‑fire segmentation in satellite imagery, comprising 2,148 image‑mask pairs from 25 California wildfires captured between July 2020 and August 2026. Each 512×512 pixel, three‑channel image is a Sentinel‑2 Level‑2A composite of bands B12, B11, and B8A, with a fixed linear rendering applied uniformly. Masks distinguish background, short‑wave‑infrared rule‑based active fire, and invalid observations, and the dataset includes chip‑level metadata, an incident‑disjoint split, and a mask‑blind analyst review of 233 test chips.

By Shreyan Mitra, Mohammadreza Narimani, Parastoo Farajpoor