The paper challenges the common practice of estimating aleatoric uncertainty in large language models (LLMs) by generating multiple clarified inputs and comparing the resulting answers. It argues that answers are unnecessary, costly, and can introduce epistemic leakage, proposing instead a clarification-only method that directly assesses ambiguity from the space of plausible interpretations. Experiments on three benchmarks show the new approach improves AUROC, reduces computational cost, and yields uncertainty estimates less correlated with epistemic uncertainty.
By Omer Nahum, Niv Nayman, Jonathan Fhima, Alon Zolfi, Jeremy Levy, Shai Mazor, Paolo Favaro
arXiv:2603. 29466v2 Announce Type: replace-cross Abstract: Existing methods for quantifying predictive uncertainty in neural networks are either computationally intractable for large language models or require access to training data that is typically unavailable.
By Nils Gr\"unefeld, Jes Frellsen, Christian Hardmeier
arXiv:2604.08974v2 Announce Type: replace
Abstract: Uncertainty quantification techniques measure confidence in language model outputs to support critical applications like hallucination detection an...
By Lorenzo Jaime Yu Flores, Cesare Spinoso di-Piano, Jackie Chi Kit Cheung
The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.
By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertaint...
arXiv:2605.25831v2 Announce Type: replace-cross
Abstract: Large language models (LLMs) define a distribution over text, which can be viewed as a probabilistic representation of uncertainty: sampling...
By Joris Baan, Wilker Aziz, Barbara Plank, Raquel Fern\'andez
arXiv:2606. 03846v1 Announce Type: cross Abstract: Large language models (LLMs) demonstrate remarkable performance across diverse tasks, but they often generate responses that appear plausible while being factually incorrect.
By Qi Cao, Takeshi Kojima, Andrew Gambardella, Helinyi Peng, Yutaka Matsuo, Yusuke Iwasawa
arXiv:2609.37594v1 Announce Type: new
Abstract: Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot t...
By David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin
The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.
By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
The paper introduces Label-Confidence-Aware Uncertainty Quantification (LCA-UQ), a method that uses Pointwise Kullback-Leibler divergence to align global entropy from multiple stochastic samples with the local confidence of a candidate answer. By bridging this gap, LCA-UQ improves the reliability and stability of uncertainty assessments in natural language generation. Experiments on popular LLMs and NLP datasets show that label sources significantly influence classification and that LCA-UQ outperforms existing uncertainty estimation approaches.
By Qinhong Lin, Yinglun Feng, Yuhao Zhang, Zhongliang Yang, Linna Zhou
Semantic Self-Distillation (SSD) is a method that distills the semantic dispersion of sampled answers from large language models into lightweight student models. These students estimate a prompt-conditioned density before answer generation, providing a prompt-level uncertainty signal via entropy and an answer-level reliability measure through probability density. Experiments on TriviaQA and MMLU show that SSD matches the teacher’s uncertainty estimates while enabling additional tasks such as hallucination prediction, out-of-domain detection, and multiple-choice answer selection.
By Edward Phillips, Sean Wu, Fredrik K. Gustafsson, Boyan Gao, David A. Clifton