arXiv Computer Vision
Sep 23

Annual Earth-observation embeddings encode wildfire disturbance and support simplified burned area mapping

Annual Earth‑observation embeddings, specifically Tessera, can encode wildfire disturbance signals well enough to map burned areas without needing curated fire‑specific imagery or dense time‑series analysis. In tests, linear models using a single Tessera embedding matched or outperformed paired pre‑ and post‑fire HLS imagery and post‑fire imagery alone, achieving high F1 scores for burn‑scar delineation and regional mapping. The approach successfully mapped all same‑year fires in benchmark scenes, recovered 97% of California burned area without California training data, transferred to European fires with high accuracy, and even estimated ignition timing within a 13‑day error margin. "whyItMatters":"The study demonstrates that pre‑trained annual embeddings can simplify and scale burned‑area mapping, reducing reliance on dense time‑series data and enabling more efficient wildfire monitoring."

By Jovana Knezevic, Clement Atzberger, Zhengpeng Feng, Adam F. A. Pellegrini, Srinivasan Keshav, David Coomes
arXiv Machine Learning
Sep 16

A Sentinel-2 benchmark dataset for deep-learning active-fire segmentation across 25 California wildfires

The article introduces an open image dataset for active‑fire segmentation in satellite imagery, comprising 2,148 image‑mask pairs from 25 California wildfires captured between July 2020 and August 2026. Each 512×512 pixel, three‑channel image is a Sentinel‑2 Level‑2A composite of bands B12, B11, and B8A, with a fixed linear rendering applied uniformly. Masks distinguish background, short‑wave‑infrared rule‑based active fire, and invalid observations, and the dataset includes chip‑level metadata, an incident‑disjoint split, and a mask‑blind analyst review of 233 test chips.

By Shreyan Mitra, Mohammadreza Narimani, Parastoo Farajpoor
arXiv Machine Learning
Sep 22

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

WildfireSpreadBench evaluates machine‑learning models for predicting next‑day wildfire spread, comparing five discriminative and one generative architecture on the WildfireSpreadTS dataset. The study shows that model rankings differ markedly when using Average Precision versus threshold‑dependent metrics such as F1 and IoU, revealing three distinct prediction profiles—over‑predicting, balanced, and under‑predicting—that AP alone cannot distinguish. Expanding input channels modestly affected AP, underscoring that AP may favor models with predictions poorly suited for operational use.

By Arin Gopakumar, Marco Pannozzo
arXiv AI
Sep 1

Toward Generalizable Deep Learning Based Peatland Fire Detection via Walsh Hadamard Transform and Domain Adaptation

arXiv:2603.02465v2 Announce Type: replace-cross Abstract: Machine learning-based wildfire detection has advanced significantly using deep learning models trained on large wildfire image and video dat...

By Emadeldeen Hamdan, Ahmad Faiz Tharima, Mohd Zahirasri Mohd Tohir, Dayang Nur Sakinah Musa, Erdem Koyuncu, Adam J. Watts, Ahmet Enis Cetin
arXiv Machine Learning
Aug 7

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

arXiv:2608. 05265v1 Announce Type: new Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas.

By Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
arXiv Machine Learning
Sep 17

Modular Deep Learning Mechanisms for Auditable Next-Day Wildfire Spread Prediction

The paper presents modular deep learning augmentations for next‑day wildfire spread prediction, including wind‑ and slope‑conditioned attention biases, physics‑feature retrieval‑augmented output correction, and fire‑conditioned dual‑stream gating. These modules are evaluated on five backbone models using the Next Day Wildfire Spread benchmark, with staged ablations, directional audits, retrieval perturbations, calibration measures, and computational comparisons. The best augmented SwinUNETR model achieves an F1 score of 0.4216 and an AUC‑PR of 0.3673, while a mixed ensemble reaches 0.4292 and 0.3790, demonstrating that predictive performance, operational trustworthiness, and computational practicality can be simultaneously improved.

By Miguel Esparza, Aydin Ayanzadeh Ahmad Mousavi, Ali Mostafavi