The paper introduces rankECE, a new metric for assessing calibration error in predictive models. Unlike the widely used Expected Calibration Error (ECE), rankECE compares predictions with neighboring probability values, offering theoretical guarantees and empirical evidence that it better approximates ECE than traditional binned methods.
By Anirban Chatterjee, Rina Foygel Barber
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
By Arthur Hoarau
arXiv:2602. 13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty.
By \'Ad\'am Jung, Domokos M. Kelen, Andr\'as A. Bencz\'ur
arXiv:2411.02771v3 Announce Type: replace-cross
Abstract: Doubly robust estimators are widely used for estimating average treatment effects and other linear summaries of regression functions. While c...
By Lars van der Laan, Alex Luedtke, Marco Carone
arXiv:2608. 07827v1 Announce Type: new Abstract: Confidence estimation for large language models (LLMs) aims to estimate the probability that a generated answer is correct, while calibration aligns these estimates with empirical accuracy.
By Avery Ma, Lorne Schell, Vin Bhaskara, Leila Pishdad
The paper argues that traditional global calibration metrics, such as Expected Calibration Error and Brier Score, are confounded by differences in model accuracy when comparing large language models. It introduces ACE, an accuracy‑controlled evaluation framework that offers Instance‑Aligned, Distribution‑Aligned, and Candidate‑Aligned views to provide fairer cross‑model comparisons. Experiments across various benchmarks reveal that many reported calibration advantages disappear after accuracy control and that model rankings often reverse, indicating that raw global metrics are unreliable for cross‑model calibration assessment.
By Zhichao Yang, Caiqi Zhang, Ruihan Yang, Chengzu Li, Nigel Collier, Deqing Yang
arXiv:2607. 27301v1 Announce Type: cross Abstract: Isotonic regression is a canonical tool for estimating monotone functions and calibrating probabilistic predictors.
By Raphael Rossellini, Rina Foygel Barber, Zhimei Ren, Jake A. Soloff
arXiv:2602.13540v2 Announce Type: replace-cross
Abstract: Accurate confidence estimation is critical for reliable use of large language models (LLMs). Prior work on LLM calibration largely focuses on...
By Sin-Han Yang, Cheng-Kuang Wu, Chieh-Yen Lin, Yun-Nung Chen, Hung-yi Lee, Shao-Hua Sun
arXiv:2511. 21140v4 Announce Type: replace Abstract: Large language models (LLMs) are widely used as scalable evaluators of model responses in lieu of human annotators.
By Chungpa Lee, Thomas Zeng, Jongwon Jeong, Jy-yong Sohn, Kangwook Lee
arXiv:2605. 30188v2 Announce Type: replace-cross Abstract: Reliable probability estimates are critical in many machine learning applications, yet modern classifiers are often poorly calibrated.
By Eug\`ene Berta, David Holzm\"uller, Francis Bach, Michael I. Jordan
arXiv:2606. 07822v1 Announce Type: cross Abstract: As language models improve and become increasingly deployed to solve a variety of tasks, trustworthiness becomes essential.
By Nishant Subramani, Palash Goyal, Yiwen Song, Mani Malek, Yuan Xue, Tomas Pfister, Hamid Palangi
arXiv:2601. 07965v2 Announce Type: replace Abstract: When a model knows when it does not know, many possibilities emerge.
By Chenjie Hao, Weyl Lu, Yuko Ishiwaka, Zengyi Li, Weier Wan, Yubei Chen