Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.
arXiv:2605. 27752v3 Announce Type: replace Abstract: Is verbalized confidence better calibrated than token likelihood?
arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.
arXiv:2606. 24281v1 Announce Type: cross Abstract: Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success.
The paper investigates whether confidence signals from fine‑tuned large language models can improve extractive question answering that relies heavily on retrieval. Experiments on four 7‑9B model families show that retrieval alone recovers 92–99.8% of the best possible accuracy, leaving little room for confidence‑based routing or adaptation to help. The sequence‑likelihood confidence metric, even after recalibration or temperature scaling, fails to provide a statistically significant benefit across different correctness criteria and answer lengths, and the study ultimately offers a set of pre‑specified negatives with explicit dependencies as its main contribution.
arXiv:2603.17839v4 Announce Type: replace-cross Abstract: Verbal confidence -- prompting LLMs to state their confidence as a number or category -- is widely used to extract uncertainty estimates from...
The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.
arXiv:2607. 20526v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly deployed in settings where fluent but incorrect answers can be costly.
The paper introduces a joint fact‑verification score that evaluates both answers and the evidence submitted with them. On the FEVEROUS dataset, replacing the DCUF evidence with UnifEE evidence improves the strict score by about 9.6 percentage points, while answer accuracy rises only 1.96 points. The study also shows that increasing context length for large language models yields modest evidence‑gain improvements, and that detailed answer‑evidence analyses uncover patterns missed by aggregate metrics.
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
arXiv:2609.33142v2 Announce Type: replace Abstract: Generating a correct answer does not mean that a language model will select it. We separate factual recall into three steps: generating a correct c...
The paper compares two common ways of evaluating large language models (LLMs): prompting them to answer questions directly and scoring candidate answers using likelihood-based metrics. The authors introduce a new protocol that ranks declarative statements derived from question–answer pairs, and test it across 95 decoder-only models (0.1B–104B parameters) on 10 multiple-choice QA datasets. They find that while prompted answering accuracy improves sharply with model scale and instruction tuning, statement‑likelihood ranking accuracy stays relatively stable, indicating that the two evaluation methods probe different aspects of model behavior.
The paper introduces XConf, an experiential confidence estimator that augments a language model’s current inference with a record of its past graded episodes. By recalling similar past tasks and reflecting on past outcomes, XConf generates confidence scores without accessing logits or updating weights, achieving superior discrimination and calibration across diverse benchmarks. The method demonstrates significant gains in selective prediction, improving success rates on agent tasks by up to 8.7 points.
The paper investigates token‑level certainty as a proxy for correctness in large language models. It finds that certainty better predicts whether a model will answer a question correctly than it does whether a specific response is correct, and that certainty varies by token type and position. The authors show that using certainty early in generation to allocate responses and later to weight votes improves accuracy while dramatically cutting token cost.