arXiv Machine Learning

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

WildfireSpreadBench evaluates machine‑learning models for predicting next‑day wildfire spread, comparing five discriminative and one generative architecture on the WildfireSpreadTS dataset. The study shows that model rankings differ markedly when using Average Precision versus threshold‑dependent metrics such as F1 and IoU, revealing three distinct prediction profiles—over‑predicting, balanced, and under‑predicting—that AP alone cannot distinguish. Expanding input channels modestly affected AP, underscoring that AP may favor models with predictions poorly suited for operational use.

arXiv Machine Learning
6d ago

Modular Deep Learning Mechanisms for Auditable Next-Day Wildfire Spread Prediction

The paper presents modular deep learning augmentations for next‑day wildfire spread prediction, including wind‑ and slope‑conditioned attention biases, physics‑feature retrieval‑augmented output correction, and fire‑conditioned dual‑stream gating. These modules are evaluated on five backbone models using the Next Day Wildfire Spread benchmark, with staged ablations, directional audits, retrieval perturbations, calibration measures, and computational comparisons. The best augmented SwinUNETR model achieves an F1 score of 0.4216 and an AUC‑PR of 0.3673, while a mixed ensemble reaches 0.4292 and 0.3790, demonstrating that predictive performance, operational trustworthiness, and computational practicality can be simultaneously improved.

By Miguel Esparza, Aydin Ayanzadeh Ahmad Mousavi, Ali Mostafavi
arXiv Machine Learning
Aug 7

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

arXiv:2608. 05265v1 Announce Type: new Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas.

By Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
arXiv AI
Aug 6

Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals

arXiv:2607. 21597v2 Announce Type: replace Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal.

By Nicolas Caron, Christophe Guyeux, Hassan Noura, Maxime Coulmeau, Benjamin Aynes
arXiv AI
6d ago

Interpretable Patch-Based Deep Learning for Wildfire Spread Prediction from Ensemble Simulations

The study evaluates deep learning surrogates for wildfire spread prediction, training four architectures on 10,584 high‑resolution simulations from Catalonia. Results show that only surface fuel load significantly predicts burn probability, and convolutional models mainly use distance to the fire front while a transformer model emphasizes fuel and terrain. When applied to a new region without retraining, the models still perform reasonably, with only a modest accuracy drop.

By Marcin Lawenda, Aleksandra Krasicka, David Caballero, Luis Torres, {\L}ukasz Szustak
arXiv Machine Learning
Sep 7

Advancing Subseasonal Forecasting with Machine Learning

The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Jul 13

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Sep 7

Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks

The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.

By Kornelius Raeth, Nicole Ludwig