As the chain-of-thought reasoning capabilities of large language models improve, evaluating and calibrating their reasoning confidence is becoming increasingly important for quantifying the uncertaint...
The paper introduces Divergent Token Confidence (DTC), a method that estimates large language model confidence by counting tokens where two models strongly disagree during decoding. DTC uses Jensen-Shannon divergence between next-token distributions along the same reasoning trajectory and shows a near-negative correlation with answer accuracy. Experiments on multiple model families and six mathematical benchmarks demonstrate that DTC improves calibration over traditional probability-based and verbalized baselines, achieving lower expected calibration errors in both white-box and black-box settings.
By Feiyang Li, Shengjing Liu, Qi Zhan, Sijie Cheng, Weiqing Wang, Hongwen Chen, Yuxuan Yang, Wen Wang, Yile Wang, Hui Huang
arXiv:2609.37594v1 Announce Type: new
Abstract: Uncertainty estimates tell us how unsure a model is, but not why. Without knowing which parts of an input influences a model's uncertainty, we cannot t...
By David Achara, Maryam Sultana, Alexander D. Rast, Fabio Cuzzolin
The paper introduces a signed lexical gate that combines a sentence classifier’s logit margin with a sparse lexical model’s support for the predicted intent, assigning positive evidence to lexical agreement and negative evidence to a lexically favored competing intent. This gate retains more information than unsigned lexical confidence or a hard agreement rule and is calibrated via an independent binomial procedure to meet specified risk targets. Experiments on BANKING77, CLINC150, and HWU64 show that the proposed score reduces the area under the risk‑coverage curve by up to 15.8% and increases accepted coverage at low error rates, offering a compact, interpretable confidence enhancement for risk‑calibrated intent routing.
By Yezhou Cheng, Zehua Yang, Bojun Lin
arXiv:2601.19918v2 Announce Type: replace
Abstract: Hallucinations in Large Language Models (LLMs), i.e., plausible but non-factual generations, pose a significant challenge to reliable deployment in...
By Yitong Qiao, Licheng Pan, Yu Mi, Lei Liu, Yue Shen, Jian Wang, Jinjie Gu, Fei Sun, Zhixuan Chu
Recent advancements in Large Language Models (LLMs) have enabled sophisticated reasoning and content generation, yet their inherent stochasticity poses significant challenges for ensuring predictive credibility. While traditional uncertainty taxonomy paradigms, such as the dichotomy of aleatoric and epistemic uncertainties, provide conceptual foundations, they often fail to capture the multi-component and multi-stage nature of LLM generation and struggle to evaluate the effectiveness of various Uncertainty Quantification (UQ) methods.
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.
By Jianru Shen
arXiv:2609.32964v2 Announce Type: replace
Abstract: Language models (LMs) often hallucinate by committing to confident answers rather than abstaining, even when they do not have enough information to...
By Vy Nguyen, Ziqi Xu, Jeffrey Chan, Estrid He, Feng Xia, Renqiang Luo, Erik Cambria, Xiuzhen Zhang
arXiv:2606. 07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential.
By Nishant Subramani, Palash Goyal, Yiwen Song, Mani Malek, Yuan Xue, Tomas Pfister, Hamid Palangi
arXiv:2607. 20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.
By Matthew ffrench-Constant, Daniel Yang, Xinmeng Huang, Sanyam Kapoor
Manacá-1B is a 1.72‑billion‑parameter, open decoder‑only language model trained from scratch for Brazilian Portuguese, released with a fully containerized, reproducible training pipeline and complete logs. The authors evaluate it against nine open baselines on four Portuguese benchmarks, reporting standard errors and paired significance tests, and find that Manacá-1B outperforms smaller models on LAMBADA‑PT while remaining competitive on commonsense completion. They also uncover a tokenizer‑related evaluation pitfall that can drastically lower accuracy and provide a simple fix, releasing all code, logs, and corrected tokenizer for full reproducibility.
By Bruno Leonardo Santos Menezes, Carlos Leonardo Souza Cardoso, Fabio Andre Machado Porto
The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.
By Yunfan Zhou, Ye Zhu, Zhihai Wang, Jianguo Yao, Haibing Guan, Xijun Li