The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.
By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
The paper presents Laxmi, a retrained version of the AIFS weather model that uses satellite-based precipitation observations instead of ERA5 reanalysis data. Laxmi achieves a 19% improvement in global probabilistic accuracy, reduces drizzle overprediction by 33%, and boosts the 95th percentile Brier skill score by 57%. In a case study of 10 Indian tropical storms, Laxmi accurately forecasted 150 mm event-total precipitation in 7 events, outperforming both the original AIFS and the leading physical model IFS.
By Julian F. Schmitt, Bertrand Delorme, Robert C. King, Yashica Patodia, Tapio Schneider, Aditi Sheshadri, Ravi Jain
Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challenges absent at shorter horizons. Errors accumulate over long autoregressive rollouts, systematic biases grow with lead time, and several years of data must be held out for independent verification, even though machine-learning models otherwise benefit from longer training records.
arXiv:2607. 05100v1 Announce Type: cross Abstract: Data-driven models now rival numerical weather prediction in the medium range, but extending them to sub-seasonal lead times raises challenges absent at shorter horizons.
By Jakob Schloer, Steffen Tietsche, Christopher D. Roberts, Lorenzo Zampieri, Simon Lang, Gert Mertes, Gareth Jones, Matthew Chantry, Frederic Vitart
arXiv:2602. 16579v2 Announce Type: replace-cross Abstract: Reliable global streamflow forecasting is essential for flood preparedness and water resource management, yet data-driven models often suffer from a performance gap when transitioning from historical reanalysis to operational forecast products.
By Maria Luisa Taccari, Kenza Tazi, Ois\'in M. Morrison, Andreas Grafberger, Juan Colonese, Corentin Carton de Wiart, Christel Prudhomme, Cinzia Mazzetti, Matthew Chantry, Florian Pappenberger
We introduce a stretched-grid artificial intelligence (AI) weather forecasting model with 2-km resolution over the western United States and part of the Northeast Pacific and approximately 31-km resol...
arXiv:2608.30795v1 Announce Type: cross
Abstract: End-to-end weather forecasting systems produce skillful global gridded and station forecasts directly from raw Earth observations, replacing the nume...
By Rodrigo Almeida, Noelia Otero, Jost Arndt, Simon Baur, Wojciech Samek, Jackie Ma
The paper introduces Nested‑EAGLE, a 0.25° global weather model with a 6 km refinement over the contiguous United States, designed to merge short‑ and medium‑range forecasts into a single system. It shows lower mean‑squared error for near‑surface and low‑level variables over the U.S. compared to NOAA’s GFS and HRRR, while remaining competitive globally. Although precipitation forecasts are less skillful than HRRR’s deterministic training, Nested‑EAGLE delivers the most accurate storm‑location predictions at longer lead times, with blurred extrema.
By Timothy A. Smith, Mariah Pope, Sergey Frolov, Brett Basarab, Daniel Abdi, Paul Madden, Isidora Jankov
The study presents a deployment‑aware framework for forecasting spring discharge and groundwater levels in the Edwards Aquifer over 1‑12 week horizons using 79 years of hydroclimatic data. Five machine‑learning families—extreme gradient boosting, extremely randomized trees, LSTM, CNN, and Transformers—were compared, with extreme gradient boosting consistently delivering the highest reliability (R² ≥ 0.94) and strong agreement with operational drought thresholds. The validated models were integrated into a five‑agent operational architecture that automates data acquisition, model selection, prediction, threshold monitoring, verification, literature retrieval, and reporting.
By Pramod Lekhak, Chetan Sharma, Hakan Ba\c{s}a\u{g}ao\u{g}lu, F. Paul Bertetti, Debaditya Chakraborty
The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
By Kornelius Raeth, Nicole Ludwig
arXiv:2509. 25017v2 Announce Type: replace Abstract: Wildfires are among the most severe natural hazards, posing a significant threat to both humans and natural ecosystems.
By Spyros Kondylatos, Nikolas Papadopoulos, Gustau Camps-Valls, Ioannis Papoutsis
arXiv:2608. 09948v1 Announce Type: cross Abstract: No single AI weather model excels at all variables, pressure levels, and lead times.
By Qiang Wu, Han Li, Jianping Huang