arXiv AI By Annika Schneider, Tommy Rochussen, Joshua Stiller, Vincent Fortuin

Decision-Aligned Evaluation of Uncertainty Quantification

Read the original on arXiv AI →

arXiv:2606. 26990v1 Announce Type: cross Abstract: Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 7

An Axiomatic Assessment of Entropy- and Variance-based Uncertainty Quantification in Regression

arXiv:2504. 18433v3 Announce Type: replace Abstract: Uncertainty quantification is crucial in machine learning, yet most (axiomatic) studies of uncertainty measures focus on classification, leaving a gap in regression settings with limited formal justification and evaluations.

By Christopher B\"ulte, Yusuf Sale, Timo L\"ohr, Paul Hofman, Gitta Kutyniok, Eyke H\"ullermeier
arXiv Machine Learning
Sep 14

Statistical Uncertainty Quantification for Aggregate Performance Metrics in Machine Learning Benchmarks

The paper discusses how to quantify statistical uncertainty for aggregate performance metrics in machine learning benchmarks, focusing on methods such as bootstrapping, Bayesian hierarchical modeling, and visualizing task weightings with standard errors. It demonstrates that these techniques can uncover insights—for example, revealing that a model may dominate specific task types even if its overall performance is poor. The authors apply their approach to the Visual Task Adaptation Benchmark (VTAB) to illustrate its practical usefulness.

By Rachel Longjohn, Giri Gopalan, Emily Casleton
arXiv Machine Learning
Sep 22

Conformal Robustness in Prediction-Driven Decision-Making

The paper introduces a score‑calibrated robustness framework that transforms any fixed point predictor into a decision‑relevant uncertainty representation using distribution‑free conformal calibration. By employing the conformal score as the core unit of robustness, the authors derive both reliability‑based robust optimization and target‑oriented Conformal Robust Satisficing formulations, linking them through a shared robust decision frontier and a fragility measure. Experiments on synthetic data and a real online‑grocery inventory case study demonstrate the framework’s ability to improve reliability, reduce costs, and provide interpretable uncertainty scales for black‑box predictors.

By Lingjie Zhao, Hansheng Jiang, Wei Qi
arXiv AI
Sep 12

Calibration-Aware Uncertainty Cascades for Efficient Heterogeneous Model Collaboration

The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.

By Yilin Zhang, Han Jiang, Cai Xu, Ying Liu, Wei Zhao