Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration
arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.
arXiv:2606. 09876v1 Announce Type: new Abstract: Large language models often express high confidence in answers that are wrong.
arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.
arXiv:2608. 13591v1 Announce Type: new Abstract: High-confidence errors in large language models are often treated as evidence of fragile internal inference.
arXiv:2508. 14390v2 Announce Type: replace-cross Abstract: Large language models (LLMs) often express verbal confidence that is poorly aligned with actual correctness, limiting their reliability in safety-critical applications.
arXiv:2606. 05799v1 Announce Type: new Abstract: Existing calibration methods for Large Language Models (LLMs) often overlook a critical dimension of trustworthiness: a model's {\em behavioral robustness} to irrelevant or misleading information.
arXiv:2606. 24281v1 Announce Type: cross Abstract: Reasoning language models are increasingly asked not only to answer difficult questions, but also to estimate their likelihood of success.
arXiv:2606. 11211v1 Announce Type: cross Abstract: The ability of large language models (LLMs) to express calibrated uncertainty is important for safe deployment.
arXiv:2608. 09080v1 Announce Type: cross Abstract: Large Language Models (LLMs) have achieved strong performance in medical question answering and clinical reasoning tasks.
arXiv:2607. 08456v1 Announce Type: cross Abstract: A model should refuse two different things: answers it would get wrong, and questions it should not answer at all, such as unanswerable ones or ones resting on a false premise.
arXiv:2606. 07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential.
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.
arXiv:2607. 07626v1 Announce Type: cross Abstract: Reliable confidence estimation is essential for deploying large language models (LLMs) in confidence-aware systems, where downstream decisions such as retrieval, tool use, and adaptive computation depend on accurately estimating answer reliability.
arXiv:2606. 05976v2 Announce Type: replace Abstract: Recent works show that LLM agents struggle to correct errors in their own reasoning traces, despite their ability to correct errors from external sources.