The paper investigates the coherence of probabilistic forecasts produced by language models, particularly when users rely on them for life decisions involving uncertain events. Using a de Finetti-based method, the authors extract forecasts from language models about stock‑return events and compute the maximum Dutch‑book profit via linear programming, a metric of incoherence that does not require observed outcomes. The study finds significant incoherence, especially when events have complex logical relationships or when irrelevant context is present, and suggests that alternative training strategies could improve probabilistic coherence.
The paper proposes measuring a language model’s understanding via no‑arbitrage, defining it as the inability of a bounded trader to profit from Dutch books against the model’s probabilities on logically related claims. It shows that full logical coherence is computationally infeasible, that standard next‑token training yields incoherent predictions across formats, and that uncertainty grows predictably along reasoning chains, creating arbitrage opportunities. The authors introduce Arbitr, a training framework that penalizes logical inconsistencies while maintaining accuracy, reducing exploitability by orders of magnitude and revealing a scaling illusion where large models appear coherent yet exhibit extreme unjustified confidence.
By Daniel Dragonevskiy
arXiv:2608.23058v1 Announce Type: new
Abstract: Large language models (LLMs) now support forecasting systems that combine language-based reasoning with temporal data, evidence retrieval, external too...
By Xiaogang Xu, Jiaqi Tang, Jianmin Chen, Yingying Yan, Zhenchao Tang, Xiangxin Zhou, Xiaobin Hu, Wei Wei, Jinfeng Wu, Qifeng Chen, Lu Zhou, Jiafei Wu, Zhe Liu, Jianwei Yin, Weimin Zheng
arXiv:2607. 20441v1 Announce Type: cross Abstract: Every information ecosystem produces beliefs that shape strategic decisions.
By Mykola Khandoga, Yevhen Kostiuk, Anton Polishko, Yurii Filipchuk, Kostiantyn Kozlov, Dmytro Zamriy, Artur Kiulian
LEAP (Likelihood Elicitation and Aggregation for Probabilistic forecasting) is a new approach that reorganizes how evidence is used in LLM-based forecasting systems. Instead of a monolithic prediction that aggregates all evidence at once, LEAP examines each evidence item separately, elicits likelihood parameters, and combines them with an explicit prior to produce a posterior distribution. The method supports continuous, single-choice, and multi-choice forecasts and has been shown to improve prediction and calibration metrics across models on a benchmark covering forecasting, information-seeking, and browsing tasks.
By Yufei Chen, Yiran Zhao, Xiaogang Xu, Qipeng Xie, Jiafei Wu, Zhe Liu
arXiv:2607. 19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function.
By Krish Matta, Atharv Naphade, Andy Zou
arXiv:2609.05905v1 Announce Type: cross
Abstract: LLM agents are increasingly used for live forecasting, where they retrieve up-to-date information and produce estimates for unresolved future events....
By Yuanpu Cao, Yongkang Du, Yurui Chang, Lu Lin, Jinghui Chen
The paper demonstrates that large language models (LLMs) used for forecasting real‑world events can be manipulated by simply publishing new articles, even without direct access to the model or its retriever. By injecting a small number of targeted news pieces into a common crawl corpus, an adversary can flip over half of the forecast probabilities and significantly degrade forecast accuracy. The study also shows that common defense strategies can be cheaply bypassed, highlighting the vulnerability of probabilistic LLM judgments to information‑supply‑chain attacks.
By Yuan Lu, Yukuan Zhang
arXiv:2407. 00890v5 Announce Type: replace-cross Abstract: This paper presents a comparative analysis evaluating the accuracy of Large Language Models (LLMs) against traditional macro time series forecasting approaches.
By Andrea Carriero, Davide Pettenuzzo, Shubhranshu Shekhar
The article argues that large language models (LLMs) are inadequate for consequential quantitative tasks such as pricing, risk assessment, and medical triage because language is a lossy representation of quantitative data that cannot be reversed. It formalizes this limitation as a property of the training representation rather than model capacity and identifies three essential properties—reproducibility, traceable lineage to source records, and calibrated uncertainty—that language substrates cannot provide. The authors propose a new class of models, Large Quantitative Models (LQMs), designed to meet these requirements.
By Reuben Vandeventer, David Imrem, David J. Wild
arXiv:2609.36914v1 Announce Type: new
Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...
By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv:2609.38545v1 Announce Type: cross
Abstract: Machine-learning signals built from financial text treat what institutions say, and what the media repeat, as evidence about value. But whoever shape...
By Ali Atiah Alzahrani