arXiv:2608. 15871v1 Announce Type: cross Abstract: Large language models (LLMs) trained on large-scale internet corpora encode extensive statistical regularities about social identities, attitudes, and political behaviour.
By Roman Neruda, Martin Bako\v{s}, Josef \v{S}lerka, V\'it Tu\v{c}ek, Petra Vidnerov\'a, Gabriela Kadlecov\'a
Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding. We study how to improve this recipe through ensembling: given a fixed number of samples, which off-the-shelf model forecasts should be combined to maximize accuracy?
The paper argues that traditional weather forecast evaluations, which focus on statistical comparisons between forecasts and observations, do not adequately capture how forecasts influence real-world decisions. It introduces decision calibration, a framework that assesses probabilistic forecast performance from the decision-maker’s perspective. Using this framework, the authors compare a machine learning model to a classical numerical weather prediction model across various weather-dependent decision tasks, finding that forecast-level performance does not reliably predict decision-level outcomes and that model rankings can shift depending on the decision context.
By Kornelius Raeth, Nicole Ludwig
arXiv:2601. 20771v2 Announce Type: replace-cross Abstract: Accurate forecasting of infectious disease incidence is critical for public health planning and timely intervention.
By Zacharias Komodromos, Kleanthis Malialis, Artemis Kontou, Panayiotis Kolios
The paper introduces a derandomization framework for stochastic majority vote classifiers, converting PAC‑Bayesian guarantees into deterministic majority vote guarantees. By applying disintegrated PAC‑Bayesian theory to the space of vote weight vectors, the authors derive two families of high‑probability generalization bounds for both data‑independent and data‑dependent ensembles. These bounds naturally lead to a self‑bounding learning algorithm that optimizes deterministic majority vote performance.
By Julien Bastian (LabHC), Benjamin Leblanc (LabHC, UJM, MALICE), Pascal Germain (LabHC, UJM, MALICE), Amaury Habrard (LabHC, UJM, MALICE), Guillaume Metzler (ERIC), Emilie Morvant (LabHC), Paul Viallard (MALT)
arXiv:2606. 29661v1 Announce Type: new Abstract: Top AI forecasting systems are approaching superforecaster-level accuracy on future world events, but still rely primarily on off-the-shelf LLMs combined with forecasting-specific context gathering and scaffolding.
By Matthew Aitchison, Scott Jeen, Toby Shevlane, Ben Day