The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
By Kornelius Raeth, Nicole Ludwig
arXiv:2608.30795v1 Announce Type: cross
Abstract: End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the nume...
By Rodrigo Almeida, Noelia Otero, Jost Arndt, Simon Baur, Wojciech Samek, Jackie Ma
arXiv:2609.38632v1 Announce Type: new
Abstract: Recent probabilistic weather forecasters train stochastic predictors with the continuous ranked probability score (CRPS) to generate each ensemble memb...
By Joonhyeong Park, Giung Nam, Hyungi Lee, Kyunghyun Cho, Byoungwoo Park, Juho Lee
The paper presents Laxmi, a retrained version of the AIFS weather model that uses satellite-based precipitation observations instead of ERA5 reanalysis data. Laxmi achieves a 19% improvement in global probabilistic accuracy, reduces drizzle overprediction by 33%, and boosts the 95th percentile Brier skill score by 57%. In a case study of 10 Indian tropical storms, Laxmi accurately forecasted 150 mm event-total precipitation in 7 events, outperforming both the original AIFS and the leading physical model IFS.
By Julian F. Schmitt, Bertrand Delorme, Robert C. King, Yashica Patodia, Tapio Schneider, Aditi Sheshadri, Ravi Jain
arXiv:2606. 26421v1 Announce Type: new Abstract: State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets.
By Cristiana Diaconu, Jonas Scholz, Aliaksandra Shysheya, Stratis Markou, Payel Mukhopadhyay, Miles Cranmer, Richard E. Turner
arXiv:2608. 09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times.
By Qiang Wu, Han Li, Jianping Huang
arXiv:2604. 09041v2 Announce Type: replace-cross Abstract: AI-based weather forecasting now rivals traditional physics-based ensembles, but state-of-the-art (SOTA) models rely on specialized architectures and massive computational budgets, creating a high barrier to entry.
By Salva R\"uhling Cachay, Duncan Watson-Parris, Rose Yu
The paper introduces Nested‑EAGLE, a 0.25° global weather model with a 6 km refinement over the contiguous United States, designed to merge short‑ and medium‑range forecasts into a single system. It shows lower mean‑squared error for near‑surface and low‑level variables over the U.S. compared to NOAA’s GFS and HRRR, while remaining competitive globally. Although precipitation forecasts are less skillful than HRRR’s deterministic training, Nested‑EAGLE delivers the most accurate storm‑location predictions at longer lead times, with blurred extrema.
By Timothy A. Smith, Mariah Pope, Sergey Frolov, Brett Basarab, Daniel Abdi, Paul Madden, Isidora Jankov
State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets. This restricts usage for under-resourced groups and severely limits fast model iteration.
The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.
By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv:2510. 09484v3 Announce Type: replace Abstract: Limited-Area Models (LAMs) enable weather forecasting over regional domains at higher resolutions than what is computationally feasible for global models.
By Erik Larsson, Joel Oskarsson, Tomas Landelius, Fredrik Lindsten
arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.
By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey