The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
By Kornelius Raeth, Nicole Ludwig
arXiv:2607. 00164v1 Announce Type: new Abstract: Reinforcement learning with verifiable rewards can in principle train calibrated probabilistic forecasters, since a proper scoring rule such as the Brier score is computed from outcomes alone and is minimized in expectation by the true probability.
By Sadanand Singh, Allam Reddy, Manan Chopra
arXiv:2606. 15917v1 Announce Type: new Abstract: We use Group Relative Policy Optimization (GRPO), a recently devised sample and memory efficient reinforcement learning method, to finetune pretrained LLMs in the range of 1.
By Amit Arnold Levy
arXiv:2401. 14483v4 Announce Type: replace Abstract: In the current practices of machine learning, the evaluation of forecasts has become a cornerstone of scientific progress.
By Rabanus Derr, Robert C. Williamson
arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.
By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?
arXiv:2609.13840v1 Announce Type: new
Abstract: A contract-logistics spare-parts operator is paid on order-level service: an order counts only if every requested line is fulfilled, yet forecasters ar...
By Joo Ern Chin, Shih-Fen Cheng, Aldy Gunawan
arXiv:2608. 03416v1 Announce Type: new Abstract: Large language models (LLMs) are now regularly asked to forecast real-world events, but comparisons are often difficult because models receive different information, use different tools, and are evaluated under different rules.
By Jonaid Shianifar, Iias Faiud
arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.
By Dean P. Foster, Sergiu Hart
The paper introduces Generalized Gibbs Ensemble Weighting (GGEW), a probabilistic framework that assigns weights to forecasting models using a Gibbs-style exponential transformation of normalized predictive loss. GGEW extends basic weighting through numerical stabilization, diversity-aware score corrections, and online hyperparameter adaptation, yielding variants such as Stable Gibbs weighting, Directional Gibbs-NCL, and Symmetric Gibbs-NCL. The authors evaluate GGEW on M4 competition submissions and real-world datasets (Monash Traffic, Electricity, Solar), finding that Gibbs-style adaptive weighting is competitive across various settings, though performance varies by dataset, horizon, and deployment protocol.
By Prasen R. Nuthanakaluva, Nava K. Gaddam
arXiv:2607. 01171v1 Announce Type: new Abstract: Sample-based generative models are increasingly used for probabilistic forecasting in high-stakes decision settings, yet their training objectives are blind to the decision maker's cost structure.
By Kornelius Raeth, Nicole Ludwig
arXiv:2608. 13554v1 Announce Type: new Abstract: We study online probabilistic forecasting of binary outcomes chosen by an adaptive adversary.
By Georgy Noarov, Aaron Roth