arXiv Machine Learning

Downside-Controlled Online Forecast Combination under Delayed and Revised Outcomes

The paper proposes a method for controlling downside risk when adjusting forecasts from frozen models, such as foundation models, by combining a static corrector and an online corrector on the simplex. Using only post‑horizon losses, the approach achieves minimal deterioration (0.15%) and up to 11.5% gains across 28 forecast pairs, and consistently reduces mean MSE in day‑ahead load forecasts for seven European bidding zones. The method’s applicability is bounded by three empirical conditions related to expert speed, stream length, and outcome alignment.

arXiv Machine Learning
4d ago

Evaluation Choices Decide the Forecasting Leaderboard: Evidence from a Production Marketplace Panel

The paper demonstrates that the outcome of a forecasting leaderboard is largely determined by the evaluator’s design choices rather than the models themselves. By fixing the data, horizon, and period, the authors varied three key evaluation decisions—unit of analysis, error pooling, and scoring metric—and showed that each can reverse or eliminate the apparent superiority of any forecasting method. The study also evaluates the practical impact of these choices on a deployed system, revealing that the selection rule captures a significant portion of the potential performance gain, and confirms the findings on an external public dataset.

By Md Rezwanul Islam, Wael Mohammed
arXiv Machine Learning
Aug 31

How Proper Scoring Rules Shape LLM Forecasting

The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.

By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
arXiv Machine Learning
Jul 13

Enhancing AI and Dynamical Subseasonal Forecasts with Probabilistic Bias Correction

arXiv:2604. 16238v2 Announce Type: replace Abstract: Decision-makers rely on weather forecasts to plant crops, manage wildfires, allocate water and energy, and prepare for weather extremes.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Sep 7

Advancing Subseasonal Forecasting with Machine Learning

The paper introduces Probabilistic Bias Correction (PBC), a machine learning framework that learns to correct historical probabilistic forecasts, thereby reducing systematic errors in subseasonal weather predictions. Applied to leading dynamical and AI models from ECMWF, PBC doubles the AI system’s modest subseasonal skill and improves the operationally-debiased dynamical model for most pressure, temperature, and precipitation targets. In ECMWF’s 2025 real‑time forecasting competition, PBC’s global forecasts ranked first across all weather variables and lead times, outperforming multiple operational and ensemble models.

By Hannah Guan, Soukayna Mouatadid, Paulo Orenstein, Judah Cohen, Haiyu Dong, Zekun Ni, Jeremy Berman, Genevieve Flaspohler, Alex Lu, Jakob Schloer, Joshua Talib, Jonathan A. Weyn, Lester Mackey
arXiv Machine Learning
Aug 19

Dynamic Regime-Aware Conformal Calibration for Reliable Economic Forecast Intervals under Multiple Distribution Shifts

Dynamic Regime-Aware Conformal Prediction (DRACP) is a new method that blends density‑ratio estimation, localized kernel weighting, and probabilistic regime‑aware weighting with a self‑tuning online significance controller to produce reliable prediction intervals under multiple distribution shifts. The authors prove finite‑sample validity with oracle weights, provide a coverage‑gap bound for estimated weights, and give deterministic or regret guarantees for the online controller. In experiments on 48 real forecasting series—including euro‑area inflation, US macroeconomic and energy indicators, and daily financial data—DRACP achieves the most reliable calibration, maintaining coverage close to the nominal 0.90 and never falling below 0.80, while other methods achieve narrower intervals but with higher under‑coverage. whyItMatters":"DRACP offers a principled trade‑off between calibration and efficiency, ensuring that prediction intervals meet coverage standards even when economic data exhibit covariate shift, concept drift, and latent regimes."

By Bogdan Oancea