arXiv AI
3d ago

Probability is Not Enough: Exploring and Counting Divergent Tokens for Reasoning Uncertainty Quantification in LLMs

The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.

By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
arXiv AI
1d ago

Signed Lexical Confidence for Risk-Calibrated Intent Routing

The paper introduces a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the predicted intent, assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent. This gate retains more information than unsigned lexical confidence or a hard agreement rule and is calibrated via an independent binomial procedure to meet specified risk targets. Experiments on BANKING77, CLINC150, and HWU64 show that the proposed score reduces the area under the risk‑coverage curve by up to 15.8% and increases accepted coverage at low error rates, offering a compact, interpretable confidence enhancement for risk‑calibrated intent routing.

By Yezhou Cheng, Zehua Yang, Bojun Lin
Hugging Face Trending Papers
Jun 22

The Origins of Stochasticity: Comprehensive Investigations on Uncertainty Quantification for Large Language Models

Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.