arXiv AI By Hankyeol Kim, Pilsung Kang

Asking Is Not Enough: Protocol Sensitivity in LLM Confidence Calibration

Read the original on arXiv AI →

arXiv:2605. 27752v2 Announce Type: replace Abstract: LLM confidence calibration is often evaluated by comparing two signals: token-probability scores and verbalized confidence.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.