arXiv:2606. 19868v1 Announce Type: new Abstract: Although large language models (LLMs) have shown strong capabilities across a wide range of tasks, their outputs often remain unreliable and may contain hallucinations, making uncertainty estimation (UE) essential for building trustworthy LLMs.
By Jiayi Wang, Xu-Yao Zhang
arXiv:2603. 24929v2 Announce Type: replace Abstract: Understanding and quantifying uncertainty in large language model (LLM) outputs is critical for reliable deployment.
By Farhan Ahmed, Yuya Jeremy Ong, Chad DeLuca
The paper introduces a method for auditing the calibration of large language models (LLMs) that only exposes a logit_bias parameter. By mathematically manipulating this parameter, the authors can evaluate exact probability thresholds with a single query per sample, enabling a provably consistent estimator of True Calibration Error for binary tasks. This approach offers an efficient framework for auditing black‑box foundation models despite limited access to continuous output probabilities.
By Roman Plaud, Antoine Saillenfest, Matthieu Labeau, Thomas Bonald, Willem Waegeman
arXiv:2604.11662v2 Announce Type: replace
Abstract: Recent work has shown that the hidden states of large language models contain signals useful for uncertainty estimation, motivating a growing inter...
By Joe Stacey, Hadas Orgad, Kentaro Inui, Benjamin Heinzerling, Nafise Sadat Moosavi
arXiv:2608. 08700v1 Announce Type: new Abstract: Reliable evaluation of tool routing is critical as Large Language Models increasingly operate as autonomous agents.
By Dongjie Xu, Julius, Hanchi Dong, Minghua Tang, Yuxuan Sun, Ziwei Nie, Zicheng Liu, Dujun Qing, Jiajie Xu
arXiv:2603. 25450v2 Announce Type: replace Abstract: Detecting when a language model is wrong without ground truth labels is a fundamental challenge for safe deployment.
By Matt Gorbett, Suman Jana