arXiv AI
Sep 1

Toward Generalizable Deep Learning Based Peatland Fire Detection via Walsh Hadamard Transform and Domain Adaptation

arXiv:2603.02465v2 Announce Type: replace-cross Abstract: Machine learning-based wildfire detection has advanced significantly using deep learning models trained on large wildfire image and video dat...

By Emadeldeen Hamdan, Ahmad Faiz Tharima, Mohd Zahirasri Mohd Tohir, Dayang Nur Sakinah Musa, Erdem Koyuncu, Adam J. Watts, Ahmet Enis Cetin
arXiv Computer Vision
Sep 2

Scale-based Approach for Active Wildfire Segmentation on Satellite Imagery

The paper proposes a data-driven protocol that uses multispectral Landsat‑8 imagery and connected‑component analysis to characterize fire‑region size distributions for active wildfire segmentation. Three segmentation architectures—U‑Net, DeepLabV3+, and SegFormer—are evaluated under different SWIR‑based spectral configurations, with U‑Net showing the strongest robustness and SWIR2 consistently delivering the best results. The study highlights the importance of both spectral band selection and architectural design for robust satellite‑based active wildfire mapping, especially when training on low fire‑pixel density images.

By Matheus F. Kovaleski, Cristiano Premebida, Jo\~ao Ruivo Paulo
arXiv Computer Vision
Sep 2

Multimodal RGB-Infrared Combination for UAV-Based Wildfire Segmentation: A Comparative Study on FLAME3

The paper examines RGB‑infrared fusion for binary wildfire segmentation using UAV imagery on the FLAME3 dataset. It compares RGB and infrared baselines with three fusion strategies across U‑Net, DeepLabV3+, and SegFormer architectures. Results show thermal data dominates segmentation performance, and feature‑level fusion with transformer‑based models yields the best results.

By Matheus F. Kovaleski, Lu\'is Garrote, Cristiano Premebida, J\'er\^ome Mendes, Jo\~ao Ruivo Paulo
arXiv AI
Sep 10

SAFIRE: Safety-Critical Benchmark for Fine-grained Fire and Smoke Understanding in Multimodal LLMs

SAFIRE is a large-scale benchmark for fire and smoke understanding in multimodal large language models (MLLMs), featuring 83,000 captioned images across 20 scenarios and 193,000 multiple-choice VQA questions derived from a 9.7K-image subset. The benchmark evaluates 10 dimensions of performance, from basic perception to higher-order reasoning, and employs a GPT‑5.4-assisted verification pipeline to ensure annotation quality. Experiments on ten open-source MLLMs (8B–38B) reveal an average accuracy of 61.9%, highlighting significant gaps in safety-critical reasoning, while fine-tuning vision encoders on just 7% of SAFIRE data boosts fire-scene classification accuracy from 20.1% to 64.5%. All resources are publicly available at https://risys-lab.github.io/SAFIRE/.

By Pengfei Li, Naufal Suryanto, Sicheng Zhang, Mohammad Alsharid, Muzammal Naseer