arXiv Machine Learning

Forecast Skill Is Not Decision Skill: Evidence from Weather-Dependent Decision Tasks

The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.

arXiv AI
Jun 3

AdaWeather: Adaptively Mixing Probabilistic Weather Forecasts with Logarithmic Regret

arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.

By Saptarishi Dhanuka (Ashoka University), Sarvesh Iyer (Ashoka University), Manmeet Singh (Western Kentucky University), Mihir More (Ashoka University), Rushil Gupta (Ashoka University), Dhruman Gupta (Ashoka University), Parthasarathi Mukhopadhyay (Ashoka University), Sandeep Juneja (Ashoka University)
arXiv Machine Learning
Sep 7

Advancing Subseasonal Forecasting with Machine Learning

The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Jul 13

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Aug 31

How Proper Scoring Rules Shape LLM Forecasting

The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.

By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
arXiv Machine Learning
Jun 26

Huracan: A skillful end-to-end data-driven system for ensemble data assimilation and weather prediction

arXiv:2508. 18486v2 Announce Type: replace-cross Abstract: Over the past few years, machine learning-based data-driven weather prediction has been transforming operational weather forecasting by providing more accurate forecasts while using a mere fraction of computing power compared to traditional numerical weather prediction (NWP).

By Zekun Ni, Jonathan Weyn, Hang Zhang, Yanfei Xiang, Jiang Bian, Weixin Jin, Kit Thambiratnam, Qi Zhang, Haiyu Dong, Hongyu Sun
arXiv Machine Learning
Aug 28

Bridging short- and medium-range weather forecasting with machine learning

The paper introduces Nested‑EAGLE, a 0.25° global weather model with a 6 km refinement over the contiguous United States, designed to merge short‑ and medium‑range forecasts into a single system. It shows lower mean‑squared error for near‑surface and low‑level variables over the U.S. compared to NOAA’s GFS and HRRR, while remaining competitive globally. Although precipitation forecasts are less skillful than HRRR’s deterministic training, Nested‑EAGLE delivers the most accurate storm‑location predictions at longer lead times, with blurred extrema.

By Timothy A. Smith, Mariah Pope, Sergey Frolov, Brett Basarab, Daniel Abdi, Paul Madden, Isidora Jankov