Credal Large Language Models for Semantic Commitment under Uncertainty
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertaint...
The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.
arXiv:2609.37594v1 Announce Type: new Abstract: Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot t...
The paper introduces a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the predicted intent, assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent. This gate retains more information than unsigned lexical confidence or a hard agreement rule and is calibrated via an independent binomial procedure to meet specified risk targets. Experiments on BANKING77, CLINC150, and HWU64 show that the proposed score reduces the area under the risk‑coverage curve by up to 15.8% and increases accepted coverage at low error rates, offering a compact, interpretable confidence enhancement for risk‑calibrated intent routing.
arXiv:2601.19918v2 Announce Type: replace Abstract: Hallucinations in Large Language Models (LLMs), i.e., plausible but non-factual generations, pose a significant challenge to reliable deployment in...
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.