arXiv Computation and Language
Sep 18

An Analysis of Training-Free Self-Reported Confidence in Language Models

The paper investigates whether language models’ self-reported confidence is meaningful without additional training. By evaluating three training‑free signals—direct verbalization, post‑hoc probability estimates, and agreement across multiple generations—on 100 TriviaQA questions, the authors find that direct verbalization alone achieves high AUROC scores (0.956 and 0.937) for correctness prediction, while agreement-based methods perform noticeably worse. Re‑eliciting confidence for the same answers shows modest score shifts and occasional decision flips, and an audit of biography claims reveals only a small confidence gap between supported and contradicted statements.

By Lukas Meyer, Sofia Rossi, Wei Chen, Thomas Laurent, Yiming Li
arXiv Computation and Language
Sep 25

Calibration Is Not Enough: Evaluating Confidence Estimation Under Language Variations

The paper introduces a new evaluation framework for confidence estimation in large language models, focusing on three properties: robustness to prompt changes, stability across semantically equivalent answers, and sensitivity to semantically different answers. It demonstrates that existing confidence estimation methods perform well on robustness and stability but often fail to detect differences in answer meaning, revealing gaps in current evaluation practices. The framework aims to guide the selection of confidence estimators for practical applications.

By Yuxi Xia, Dennis Ulmer, Terra Blevins, Yihong Liu, Hinrich Sch\"utze, Benjamin Roth