Multi-agent systems are increasingly used for forecasting future events, as deliberation among multiple LLMs is believed to improve reasoning and calibration. Yet existing approaches overlook a critical design choice: what information each agent receives.
LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting) is a new approach that reorganizes how evidence is used in LLM-based forecasting systems. Instead of a monolithic prediction that aggregates all evidence at once, LEAP examines each evidence item separately, elicits likelihood parameters, and combines them with an explicit prior to produce a posterior distribution. The method supports continuous, single-choice, and multi-choice forecasts and has been shown to improve prediction and calibration metrics across models on a benchmark covering forecasting, information-seeking, and browsing tasks.
By Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
arXiv:2609.05905v1 Announce Type: cross
Abstract: LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events....
By Yuanpu Cao, Yongkang Du, Yurui Chang, Lu Lin, Jinghui Chen
arXiv:2606. 02497v1 Announce Type: new Abstract: Time series forecasting has advanced rapidly, especially with the emergence of foundation models that show strong zero-shot performance on numerical extrapolation.
By Yuhua Liao, Zetian Wang, Qiangqiang Nie, Zhenhua Zhang
arXiv:2607. 06157v1 Announce Type: cross Abstract: Deliberation plays a crucial role in collaboration; when humans work together, they naturally engage in communication to align information and reach an agreement.
By Chenxu Wang, Yongkun Yang, Boyuan Du, Shiwei Lin, Huaping Liu
The paper investigates whether in-context learning (ICL) in large language model agents reflects genuine recursive reasoning or simply statistical extrapolation. By testing LLM agents in a public goods game with manipulated historical feedback, the authors compare decision quality to a rational expectations equilibrium benchmark. They find that disrupting historical patterns eliminates the benefits of longer context, especially in highly interdependent settings, indicating that ICL behavior aligns more with statistical extrapolation than strategic reasoning.
By Yu Liu, Wenwen Li, Yifan Dou, Guangnan Ye
arXiv:2604. 18576v4 Announce Type: replace Abstract: We present the Bayesian Linguistic Forecaster (BLF), an agentic system for binary forecasting that achieves state-of-the-art performance on the ForecastBench benchmark.
By Kevin Murphy
arXiv:2605. 25929v2 Announce Type: replace-cross Abstract: The effectiveness of multi-agent LLM deliberation depends not only on the agents' individual predictions, but also on how they communicate and collaborate.
By Franka Bause, Jonas Niederle, Martin Pawelczyk, Rebekka Burkholz
arXiv:2608. 03031v1 Announce Type: new Abstract: Time series forecasting is fundamental to decision-making in complex systems, where future dynamics are influenced not only by historical observations but also by evolving contextual features.
By Xiaoyu Tao, Mingyue Cheng, Bokai Pan, Chuang Jiang, Huanjian Zhang, Tian Gao, Yaguo Liu, Qi Liu, Enhong Chen
arXiv:2606. 11816v1 Announce Type: cross Abstract: Forecasting real-world events requires language-model agents to reason under uncertainty from incomplete, time-bounded information.
By Yizhou Chi, Eric Chamoun, Zifeng Ding, Andreas Vlachos
Evaluating reasoning quality in multi-agent LLM systems is challenging, especially for open-ended tasks without reference answers. We investigate whether intrinsic confidence signals, token-level log-probabilities from decoding, can predict reasoning quality as assessed by LLM-as-judge evaluation.
arXiv:2608. 11420v1 Announce Type: new Abstract: Medical diagnostic reasoning is a high-impact use case for LLMs that carries significant implications for the health and wellbeing of users.
By Del Coburn, Scott Sanner, Dan Silver