arXiv:2607. 19367v1 Announce Type: new Abstract: Calibration is the primary criterion for evaluating LLM confidence, but it is insufficient: it admits trivially incoherent estimators, depends on the evaluation distribution, and does not test the extent to which the estimation can be interpreted as a consistent, underlying probability function.
By Krish Matta, Atharv Naphade, Andy Zou
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertaint...
The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.
By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...
arXiv:2607. 03882v1 Announce Type: cross Abstract: LLMs are increasingly deployed as post-hoc explainers of AI-generated outputs, yet it remains unclear whether they can reliably communicate probabilistic information in natural language.
By Diego Cerda-Mardini, Sarath Chandar, Sreenath Madathil
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.
By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.
By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.
By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
The paper examines how large language models (LLMs) express uncertainty compared to humans, noting that humans use verbal markers like "possible" or "likely" to convey metacognitive awareness. By curating a corpus of human uncertainty markers and benchmarking LLMs against it, the authors find that LLMs encode these markers with numerical levels that differ substantially from human usage. They introduce METHODNAME, an optimization-based algorithm that learns an optimal uncertainty profile over verbal markers directly from LLM outputs, enabling a direct comparison of confidence semantics and revealing systematic disparities in verbal expressions.
By Jinhao Duan, Zicheng Liu, Zijie Liu, Kaidi Xu, Tianlong Chen
arXiv:2507. 06722v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations.
By Sunwoo Kim, Haneul Yoo, Alice Oh
arXiv:2608.28382v1 Announce Type: new
Abstract: Users often ask large language models (LLMs) to report how confident they are, but it is unclear whether such linguistic confidence tracks the model's...
By Hefan Zhang, Bingquan Zhang, Ming Cheng, Saeed Hassanpour, Weicheng Ma, Soroush Vosoughi