arXiv AI
Jul 7

The Anatomy of Uncertainty in LLMs

arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.

By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
arXiv Computation and Language
Sep 23

Semantic Self-Distillation for Language Model Uncertainty

Semantic Self-Distillation (SSD) is a method that distills the semantic dispersion of sampled answers from large language models into lightweight student models. These students estimate a prompt-conditioned density before answer generation, providing a prompt-level uncertainty signal via entropy and an answer-level reliability measure through probability density. Experiments on TriviaQA and MMLU show that SSD matches the teacher’s uncertainty estimates while enabling additional tasks such as hallucination prediction, out-of-domain detection, and multiple-choice answer selection.

By Edward Phillips, Sean Wu, Fredrik K. Gustafsson, Boyan Gao, David A. Clifton
arXiv AI
2d ago

Verbalized and Internal Probabilities Are Coupled in Large Language Models

The paper investigates the relationship between a large language model’s internal probability distribution and its verbalized confidence statements. By systematically manipulating training and in‑context data, the authors show that both internal and verbalized probabilities are influenced by distributional and asserted uncertainty in the data. They find that verbalized probabilities align with internal ones beyond what would be expected if they tracked the same sources independently, indicating that verbalized confidence can serve as a probe of the model’s internal distribution.

By Sinead Williamson, Jiaxuan Li, Nick Foti, Russ Webb, Masha Fedzechkina
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.