arXiv AI

Risk Is Not the Target: A Monotonic Framework for Evaluating Wildfire Operational Risk Signals

arXiv:2607. 21597v2 Announce Type: replace Abstract: Evaluating wildfire risk systems using standard machine-learning metrics such as F1-score or IoU is fundamentally flawed: these metrics assess event prediction accuracy, not the operational coherence of a continuous risk signal.

arXiv Machine Learning
Sep 22

WildfireSpreadBench: The Metric Decides the Model in Wildfire Spread Prediction

WildfireSpreadBench evaluates machine‑learning models for predicting next‑day wildfire spread, comparing five discriminative and one generative architecture on the WildfireSpreadTS dataset. The study shows that model rankings differ markedly when using Average Precision versus threshold‑dependent metrics such as F1 and IoU, revealing three distinct prediction profiles—over‑predicting, balanced, and under‑predicting—that AP alone cannot distinguish. Expanding input channels modestly affected AP, underscoring that AP may favor models with predictions poorly suited for operational use.

By Arin Gopakumar, Marco Pannozzo
arXiv Machine Learning
Sep 17

Modular Deep Learning Mechanisms for Auditable Next-Day Wildfire Spread Prediction

The paper presents modular deep learning augmentations for next‑day wildfire spread prediction, including wind‑ and slope‑conditioned attention biases, physics‑feature retrieval‑augmented output correction, and fire‑conditioned dual‑stream gating. These modules are evaluated on five backbone models using the Next Day Wildfire Spread benchmark, with staged ablations, directional audits, retrieval perturbations, calibration measures, and computational comparisons. The best augmented SwinUNETR model achieves an F1 score of 0.4216 and an AUC‑PR of 0.3673, while a mixed ensemble reaches 0.4292 and 0.3790, demonstrating that predictive performance, operational trustworthiness, and computational practicality can be simultaneously improved.

By Miguel Esparza, Aydin Ayanzadeh Ahmad Mousavi, Ali Mostafavi
arXiv Machine Learning
Sep 7

Advancing Subseasonal Forecasting with Machine Learning

The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Jul 13

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Aug 7

Evaluating Machine Learning Models for Post-Wildfire Debris-Flow Prediction

arXiv:2608. 05265v1 Announce Type: new Abstract: Prediction of post-wildfire debris flows is critical for mitigating hazards to communities, infrastructure, and resources during intense rainfall in recently burned areas.

By Quinn Ledingham, Zhengsen Xu, Yimin Zhu, Zack Dewis, Mabel Heffring, Saeid Taleghanidoozdoozan, Motasem Alkayid, Megan Greenwood, Lincoln Linlin Xu
arXiv Machine Learning
Aug 19

Proactive Road Safety Intervention in Australia: Predicting Risky Driving Hotspots from Connected Vehicle Data

The paper proposes a proactive approach to road safety in Greater Sydney by using connected vehicle telemetry to predict risky driving events before crashes occur. It quantifies risky driving with g‑force thresholds and builds spatio‑temporal heatmaps to locate high‑risk zones. Eight predictive models were compared, with ARIMA achieving the lowest error and showing that simple time‑series methods can rival deep learning when data are limited, highlighting the value of IoT data for targeted safety interventions.

By Adriana-Simona Mih\u{a}i\c{t}\u{a}, Clarence Cheung, Artur Grigorev, Tuo Mao, David Lillo-Trynes
arXiv Machine Learning
Sep 22

Predictors and Orchestrators: Parsimonious Machine Learning within an Agentic AI Harness for Multi-Horizon Karst Aquifer Forecasting

The study presents a deployment‑aware framework for forecasting spring discharge and groundwater levels in the Edwards Aquifer over 1‑12 week horizons using 79 years of hydroclimatic data. Five machine‑learning families—extreme gradient boosting, extremely randomized trees, LSTM, CNN, and Transformers—were compared, with extreme gradient boosting consistently delivering the highest reliability (R² ≥ 0.94) and strong agreement with operational drought thresholds. The validated models were integrated into a five‑agent operational architecture that automates data acquisition, model selection, prediction, threshold monitoring, verification, literature retrieval, and reporting.

By Pramod Lekhak, Chetan Sharma, Hakan Ba\c{s}a\u{g}ao\u{g}lu, F. Paul Bertetti, Debaditya Chakraborty