Temperature Scaling Is Not Enough: Calibration Gaps Under Human Label Distributions
arXiv:2607. 13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning.
arXiv:2607. 18162v1 Announce Type: new Abstract: The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al.
arXiv:2607. 13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning.
arXiv:2410.18321v3 Announce Type: replace Abstract: Confidence calibration matters wherever a classifier's probabilities, not just its labels, are consumed downstream. We study Focal Calibration Loss...
arXiv:2608. 10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining.
Signal‑Routed Temperature Scaling (SRTS‑BCE) is a 10‑parameter, argmax‑preserving calibration method that separates calibration objectives from adaptive capacity. It cross‑fits a correctness‑risk score over six logit statistics and assigns a top‑label BCE temperature to each of three risk groups, generalizing TvA‑TS when K=1. Experiments on fine‑tuned CIFAR‑100 and ViT‑B/16 show that SRTS‑BCE reduces ECE from 1.65 to 0.96 with a small calibration budget, outperforming higher‑capacity SMART+BCE when only 250 examples are available, and revealing a budget‑dependent ranking reversal on Swin‑T.
arXiv:2609.26839v1 Announce Type: cross Abstract: Post-hoc probability calibration is usually evaluated under an optimistic assumption: the held-out calibration labels are clean. In many AI deploymen...
arXiv:2609.24303v1 Announce Type: new Abstract: Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident t...
The paper introduces SMART, a lightweight recalibration technique that adjusts logits based on the margin between the top two logits, called the logit gap. It uses a soft-binned Expected Calibration Error objective to balance bias and variance, enabling stable updates even with limited calibration data. Experiments across various datasets and architectures show SMART achieves state‑of‑the‑art calibration with fewer parameters than existing methods.
arXiv:2609.38917v1 Announce Type: new Abstract: A classifier's conditional accuracy can change while its confidence distribution stays exactly the same. We study the worst-case movement of the reliab...
arXiv:2606. 03245v1 Announce Type: cross Abstract: Concepts of calibration formalize the compatibility between probabilistic predictions and the respective outcomes.
arXiv:2607. 10804v1 Announce Type: new Abstract: Graph neural networks (GNNs) are increasingly deployed in real-world applications where distribution shift is un-avoidable.
The paper introduces rankECE, a new metric for assessing calibration error in predictive models. Unlike the widely used Expected Calibration Error (ECE), rankECE compares predictions with neighboring probability values, offering theoretical guarantees and empirical evidence that it better approximates ECE than traditional binned methods.
arXiv:2608. 05064v1 Announce Type: cross Abstract: Small open-weight language models increasingly run in private, offline, and cost-sensitive settings, where the key deployment question is not only what a model answers but when it should defer to a human.