arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.
By Saptarishi Dhanuka (Ashoka University), Sarvesh Iyer (Ashoka University), Manmeet Singh (Western Kentucky University), Mihir More (Ashoka University), Rushil Gupta (Ashoka University), Dhruman Gupta (Ashoka University), Parthasarathi Mukhopadhyay (Ashoka University), Sandeep Juneja (Ashoka University)
The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.
By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.
By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.
By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
arXiv:2606. 08587v1 Announce Type: cross Abstract: Statistical post-processing has proven to be an effective tool in improving ensemble forecast of different weather variables.
By \'Agnes Baran, M\'at\'e Mihalina
arXiv:2606. 26421v1 Announce Type: new Abstract: State-of-the-art medium-range AI weather models can outperform traditional Numerical Weather Prediction (NWP) but require massive training budgets.
By Cristiana Diaconu, Jonas Scholz, Aliaksandra Shysheya, Stratis Markou, Payel Mukhopadhyay, Miles Cranmer, Richard E. Turner