A key factor in deciding whether to trust an automatic prediction is its confidence score, which should be calibrated to match the actual probability of the prediction being correct. Most confidence c...
arXiv:2607. 08299v2 Announce Type: replace Abstract: Diagnostic decision making often relies on a sequence of pathology tests that bridge patient symptoms and final disease diagnosis.
By Abu Rafe Md Jamil, Nayan Malakar
arXiv:2609.36532v1 Announce Type: cross
Abstract: In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a...
By Futoshi Futami, Jerry Huang, Ichiro Takeuchi
arXiv:2609.38705v1 Announce Type: new
Abstract: Foundation models for computational pathology are usually evaluated using AUC and accuracy, while calibration is often left untested. This matters beca...
By Wenjun Liu, Saeed Hassanpour
arXiv:2609.24303v1 Announce Type: new
Abstract: Post-training often improves task performance but can degrade confidence calibration, leaving post-trained language models (PoLMs) more overconfident t...
By Linhan Luo, Lequan Lin, Dai Shi, Feng Chen, Jos\'e Miguel Hern\'andez-Lobato, Junbin Gao
arXiv:2607. 13423v1 Announce Type: new Abstract: Temperature scaling is the dominant post-hoc calibration method in modern deep learning.
By Wisdom Dogah