arXiv AI

Decision-Aligned Evaluation of Uncertainty Quantification

arXiv:2606. 26990v1 Announce Type: cross Abstract: Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions.

arXiv Machine Learning
Jul 7

An Axiomatic Assessment of Entropy- and Variance-based Uncertainty Quantification in Regression

arXiv:2504. 18433v3 Announce Type: replace Abstract: Uncertainty quantification is crucial in machine learning, yet most (axiomatic) studies of uncertainty measures focus on classification, leaving a gap in regression settings with limited formal justification and evaluations.

By Christopher B\"ulte, Yusuf Sale, Timo L\"ohr, Paul Hofman, Gitta Kutyniok, Eyke H\"ullermeier
arXiv Machine Learning
Sep 14

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

The paper discusses how to quantify statistical uncertainty for aggregate performance metrics in machine learning benchmarks, focusing on methods such as bootstrapping, Bayesian hierarchical modeling, and visualizing task weightings with standard errors. It demonstrates that these techniques can uncover insights—for example, revealing that a model may dominate specific task types even if its overall performance is poor. The authors apply their approach to the Visual Task Adaptation Benchmark (VTAB) to illustrate its practical usefulness.

By Rachel Longjohn, Giri Gopalan, Emily Casleton
arXiv Machine Learning
Sep 22

Conformal Robustness in Prediction-Driven Decision-Making

The paper introduces a score‑calibrated robustness framework that transforms any fixed point predictor into a decision‑relevant uncertainty representation using distribution‑free conformal calibration. By employing the conformal score as the core unit of robustness, the authors derive both reliability‑based robust optimization and target‑oriented Conformal Robust Satisficing formulations, linking them through a shared robust decision frontier and a fragility measure. Experiments on synthetic data and a real online‑grocery inventory case study demonstrate the framework’s ability to improve reliability, reduce costs, and provide interpretable uncertainty scales for black‑box predictors.

By Lingjie Zhao, Hansheng Jiang, Wei Qi
arXiv AI
Sep 12

Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.

By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao
arXiv Machine Learning
Jul 30

Cost-Sensitive Conformal Prediction and Human-in-the-Loop Abstention for Imbalanced High-Stakes Decision Support: A Multi-Domain Benchmark

arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.

By Manpreet Singh, Akshatha Srikantha, Shyamal Lakhanpal
arXiv Machine Learning
Jul 21

BACON: Budgeted Human Calibration for Modeling and Evaluation with Multiple AI Judges

arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.

By Lei Shi, Anlan Zhang, Rita Lyu, Zhengmian Hu, Tong Yu, David Arbour, Avi Feller, Saayan Mitra, Ritwik Sinha