arXiv:2608. 15871v1 Announce Type: cross Abstract: Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour.
By Roman Neruda, Martin Bako\v{s}, Josef \v{S}lerka, V\'it Tu\v{c}ek, Petra Vidnerov\'a, Gabriela Kadlecov\'a
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?
The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
By Kornelius Raeth, Nicole Ludwig
arXiv:2601. 20771v2 Announce Type: replace-cross Abstract: Accurate forecasting of infectious disease incidence is critical for public health planning and timely intervention.
By Zacharias Komodromos, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios
The paper introduces a derandomization framework for stochastic majority vote classifiers, converting PAC‑Bayesian guarantees into deterministic majority vote guarantees. By applying disintegrated PAC‑Bayesian theory to the space of vote weight vectors, the authors derive two families of high‑probability generalization bounds for both data‑independent and data‑dependent ensembles. These bounds naturally lead to a self‑bounding learning algorithm that optimizes deterministic majority vote performance.
By Julien Bastian (LabHC), Benjamin Leblanc (LabHC, UJM, MALICE), Pascal Germain (LabHC, UJM, MALICE), Amaury Habrard (LabHC, UJM, MALICE), Guillaume Metzler (ERIC), Emilie Morvant (LabHC), Paul Viallard (MALT)
arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.
By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day
arXiv:2609.38196v1 Announce Type: cross
Abstract: Accurate time series forecasting is critical across various domains, yet traditional ensemble methods often suffer from the disproportionate influenc...
By Ahmad Shahi, Mamehgol Yousefi, Brendon J. Woodford, Farhaan Mirza, Tapabrata Chakraborti
arXiv:2606. 08587v1 Announce Type: cross Abstract: Statistical post-processing has proven to be an effective tool in improving ensemble forecast of different weather variables.
By \'Agnes Baran, M\'at\'e Mihalina
The paper investigates how different proper scoring rules influence the performance and behavior of large language model (LLM) forecasters. Five scoring rules were compared as training objectives for binary forecasts of real-world events, revealing that while they all theoretically incentivize truthful probability reporting, they produce models with varying calibration, probability usage, and bias, information, and noise profiles. The Brier-trained model achieved the lowest Brier score and highest AUC-ROC, whereas the log-trained model achieved the best log score and lowest calibration error, indicating that scoring rule choice can shape both forecast accuracy and error structure.
By Benjamin Turtel, Paul Wilczewski, Kris Skotheim, Ville A. Satop\"a\"a, Philip E. Tetlock
The paper introduces Generalized Gibbs Ensemble Weighting (GGEW), a probabilistic framework that assigns weights to forecasting models using a Gibbs-style exponential transformation of normalized predictive loss. GGEW extends basic weighting through numerical stabilization, diversity-aware score corrections, and online hyperparameter adaptation, yielding variants such as Stable Gibbs weighting, Directional Gibbs-NCL, and Symmetric Gibbs-NCL. The authors evaluate GGEW on M4 competition submissions and real-world datasets (Monash Traffic, Electricity, Solar), finding that Gibbs-style adaptive weighting is competitive across various settings, though performance varies by dataset, horizon, and deployment protocol.
By Prasen R. Nuthanakaluva, Nava K. Gaddam
arXiv:2606. 02663v1 Announce Type: cross Abstract: Recent advances in machine learning have produced probabilistic weather forecasting models comparable to state-of-the-art numerical weather predictors.
By Saptarishi Dhanuka (Ashoka University), Sarvesh Iyer (Ashoka University), Manmeet Singh (Western Kentucky University), Mihir More (Ashoka University), Rushil Gupta (Ashoka University), Dhruman Gupta (Ashoka University), Parthasarathi Mukhopadhyay (Ashoka University), Sandeep Juneja (Ashoka University)
The paper introduces an adaptive Mixture-of-Experts (MoE) framework for time series forecasting that incorporates expert-specific losses to give each expert a direct learning signal independent of gating weights. The overall objective combines base forecasting loss with these expert losses, encouraging experts to specialize on different temporal segments. A partial online learning strategy is added for efficient incremental updates, and experiments on economic, tourism, and energy datasets show the method outperforms state‑of‑the‑art neural models and foundation models, with ablation studies confirming the benefit of expert loss integration.
By Btissame El Mahtout, Florian Ziel