arXiv Computation and Language

Kernel Token Contradiction: a Fast and Principled Approach for LLM Claim Uncertainty Quantification

arXiv Machine Learning
Aug 27

Functional Entropy: Predicting Functional Correctness in LLM-Generated Code with Uncertainty Quantification

The paper introduces Functional Entropy, a new uncertainty quantification technique for assessing the functional correctness of code generated by large language models. It evaluates token‑probability and sampling‑based methods across three programming languages and five LLMs, finding that token‑probability approaches generalize well while NLI‑based sampling fails due to semantic clustering. Functional equivalence methods, which replace NLI with an LLM‑based functional assessment, achieve superior AUROC and calibration in most model‑benchmark combinations.

By Dylan Bouchard, Mohit Singh Chauhan, Zeya Ahmad, Ho-Kyeong Ra
arXiv AI
Sep 3

From Tokens to Semantics: Leveraging Complementary Signals for Hallucination Detection in Black-Box LLMs

The paper investigates hallucination detection in black‑box large language models by leveraging two accessible signals: semantic entropy, which captures disagreement among sampled response meanings, and token‑level uncertainty derived from log‑probabilities. It introduces a TopK aggregation technique, a hybrid CoCoA method combining uncertainty with semantic dissimilarity, and two supervised approaches—Gated and Stacked—that integrate token and semantic features. Across seven benchmarks and four language models, the supervised Stacked method performs best in many cases, while TopK and CoCoA remain competitive without labeled data, though all methods require careful threshold calibration.

By Urja Pawar, Rajitha Ramanayake, Owen O'Neill, Nabeel Kemal, Abhishek Mandal, Houssem Chatbri, Christopher Martin
arXiv AI
Jun 3

SeSE: Black-Box Uncertainty Quantification for Large Language Models Based on Structural Information Theory

arXiv:2511. 16275v4 Announce Type: replace-cross Abstract: Reliable uncertainty quantification (UQ) is essential for deploying large language models (LLMs) in safety-critical scenarios, as it enables them to abstain from responding when uncertain, thereby avoiding hallucinations, i.

By Xingtao Zhao, Hao Peng, Dingli Su, Xianghua Zeng, Chunyang Liu, Jinzhi Liao, Philip S. Yu
arXiv AI
Sep 30

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.

By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.