arXiv:2507. 06722v2 Announce Type: replace-cross Abstract: Understanding how large language models (LLMs) internally represent and process their predictions is central to detecting uncertainty and preventing hallucinations.
By Sunwoo Kim, Haneul Yoo, Alice Oh
arXiv:2509.24988v2 Announce Type: replace-cross
Abstract: Generating accurate and calibrated confidence estimates is critical for deploying LLMs in high-stakes or user-facing applications, and remain...
By Hanqi Xiao, Vaidehi Patil, Hyunji Lee, Elias Stengel-Eskin, Mohit Bansal
arXiv:2608. 05995v1 Announce Type: new Abstract: Reliable uncertainty estimates are critical in safety-sensitive applications, where understanding the sources of predictive uncertainty is essential.
By Frieder Wizgall, Georg Tirpitz, Moritz Seiler, Kerstin Ritter, B\'alint Mucs\'anyi
arXiv:2603. 24967v2 Announce Type: replace Abstract: Understanding why a large language model (LLM) is uncertain about the response is important for their reliable deployment.
By Aditya Taparia, Ransalu Senanayake, Kowshik Thopalli, Vivek Narayanaswamy
arXiv:2603. 24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment.
By Farhan Ahmed, Yuya Jeremy Ong, Chad DeLuca
arXiv:2606. 03969v1 Announce Type: cross Abstract: Reliable uncertainty communication is critical to the trustworthiness of LLMs, yet faithful calibration (FC)--the alignment between models' intrinsic and (linguistically) expressed confidence--is a persistent failure mode.
By Areeb Gani, Asal Meskin, Gabrielle Kaili-May Liu, Arman Cohan
The paper introduces a ground‑truth framework for disentangling uncertainty into epistemic and aleatoric components using sample‑conditional pointwise posterior risk. It evaluates current methods, finding that Spectral‑normalized Neural Gaussian Processes and Variational Latent Gaussian Processes best recover the ground‑truth uncertainty, while most methods align more closely with posterior variance and miss predictor bias. The study also explores the entanglement of estimated uncertainties and the impact of modeling choices, providing practical guidance and releasing 13 semi‑synthetic datasets for further validation.
By Frieder Wizgall, Georg Tirpitz, Moritz Seiler, Kerstin Ritter, B\'alint Mucs\'anyi
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
arXiv:2602.02427v3 Announce Type: replace
Abstract: Large Language Models (LLMs) have achieved significant breakthroughs across various domains, but they can still produce unreliable or misleading ou...
By Qihao Wen, Jiahao Wang, Yang Nan, Pengfei He, Ravi Tandon, Han Xu
arXiv:2608. 10430v1 Announce Type: cross Abstract: Large Language Models (LLMs) deployed as AI agents frequently exhibit user specification-grounding failures, executing hallucinated, undesired actions to force a resolution rather than expressing uncertainty.
By Sanidhya Vijayvargiya, Rahul Lokesh
arXiv:2606. 32032v1 Announce Type: cross Abstract: Metacognition is a critical component of intelligence that describes the ability to monitor and regulate one's own cognitive processes.
By Gabrielle Kaili-May Liu, Avi Caciularu, Gal Yona, Idan Szpektor, Arman Cohan
arXiv:2609.09855v1 Announce Type: new
Abstract: Although probabilistic statements are ubiquitous, foundational disagreements persist about their understanding, as exemplified by debates between Bayes...
By Benedikt H\"oltgen