arXiv:2603.17839v4 Announce Type: replace-cross
Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...
By Dharshan Kumaran, Arthur Conmy, Federico Barbero, Simon Osindero, Viorica Patraucean, Petar Veli\v{c}kovi\'c
Large language models increasingly produce and interpret verbal probability expressions, yet whether these expressions carry consistent meaning across models (or match human perceptions of uncertainty...
The study evaluates how large language models (LLMs) interpret verbal probability expressions by mapping words to numbers and testing consistency across 19 models. Results show that LLMs largely mirror human benchmarks—preserving word order, recovering key anchor points, and reflecting the high variance of the term "possible"—but they exhibit a systematic upward bias for negative expressions like "unlikely" and "improbable." Explanation elicitation reduces within‑model variance but increases divergence between models, while a bidirectional roundtrip test reveals that leading models maintain coherent internal representations.
By Christos Petridis, Konstantinos Pelechrinis, Zoran Obradovic
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.
By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertaint...