arXiv:2609.09855v1 Announce Type: new
Abstract: Although probabilistic statements are ubiquitous, foundational disagreements persist about their understanding, as exemplified by debates between Bayes...
By Benedikt H\"oltgen
arXiv:2602. 13362v2 Announce Type: replace-cross Abstract: A key challenge in probabilistic regression is ensuring that predictive distributions accurately reflect true empirical uncertainty.
By \'Ad\'am Jung, Domokos M. Kelen, Andr\'as A. Bencz\'ur
The paper introduces rankECE, a new metric for assessing calibration error in predictive models. Unlike the widely used Expected Calibration Error (ECE), rankECE compares predictions with neighboring probability values, offering theoretical guarantees and empirical evidence that it better approximates ECE than traditional binned methods.
By Anirban Chatterjee, Rina Foygel Barber
arXiv:2507. 08150v4 Announce Type: replace-cross Abstract: Accurate uncertainty quantification is critical for reliable predictive modeling.
By Ilia Azizi, Juraj Bodik, Jakob Heiss, Bin Yu
arXiv:2607. 18162v1 Announce Type: new Abstract: The soft-label Bayes-error estimator beta(z) = E[min(z, 1-z)] of Ishida et al.
By Shreyas Pradeepkumar Khandale
The paper introduces the concept of observational multiplicity, where multiple probabilistic classifiers can perform similarly yet produce conflicting predictions, undermining interpretability and safety. It proposes measuring this arbitrariness through a regret metric that captures how predictions could shift with different training labels. The authors present a general method to estimate regret, show it varies across dataset groups, and discuss its use for safety via abstention and targeted data collection.
By Erin George, Deanna Needell, Berk Ustun
arXiv:2609.36532v1 Announce Type: cross
Abstract: In multiclass probabilistic prediction, Utility Calibration (UC), which focuses auditing on specified utilities, has recently received attention as a...
By Futoshi Futami, Jerry Huang, Ichiro Takeuchi
The paper investigates the feasibility of exact truthfulness in calibration measures for sequential binary prediction. It proves that exact truthfulness cannot coexist with completeness and soundness, even when outcomes are independent. The authors then provide two reductions that transform any base calibration measure into additively or multiplicatively approximately truthful ones, achieving a multiplicative truthfulness guarantee that improves upon previous results.
By Anagha Gokul, Jason Hartline, Lunjia Hu, Jonathan Ullman, Yifan Wu
arXiv:2606. 25188v1 Announce Type: new Abstract: Efficient uncertainty quantification (UQ) is essential for trustworthy large-scale learning.
By Kun Jin, James Harrison, Jiawei Li, Sihan Liu, Jiayi Liu, Randolph Linderman, Yuening Li, Arnab Bhadury, Sourabh Prakash Bansod, Liang Liu, Jasper Snoek
arXiv:2605. 26703v2 Announce Type: replace-cross Abstract: The classic concept of "calibrated forecasts" and its more recent refinement, "calibeating," are defined with respect to the standard quadratic scoring rule.
By Dean P. Foster, Sergiu Hart
arXiv:2609.25388v1 Announce Type: cross
Abstract: A classical question in statistics is which observable quantities to condition on when drawing inferences about unobservable targets. For conformal p...
By Xuelin Yang, Baihe Huang, Yilong Hou, Guido Imbens, Michael I. Jordan
arXiv:2608. 10372v1 Announce Type: new Abstract: Post-hoc calibration aligns a classifier's predicted confidences with its empirical accuracy without retraining.
By Lening Zhao, Qipeng Zhan, Li Shen