Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning
arXiv:2607. 14817v1 Announce Type: cross Abstract: Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning.
arXiv:2606. 26990v1 Announce Type: cross Abstract: Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions.
arXiv:2607. 14817v1 Announce Type: cross Abstract: Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning.
arXiv:2504. 18433v3 Announce Type: replace Abstract: Uncertainty quantification is crucial in machine learning, yet most (axiomatic) studies of uncertainty measures focus on classification, leaving a gap in regression settings with limited formal justification and evaluations.
The paper discusses how to quantify statistical uncertainty for aggregate performance metrics in machine learning benchmarks, focusing on methods such as bootstrapping, Bayesian hierarchical modeling, and visualizing task weightings with standard errors. It demonstrates that these techniques can uncover insights—for example, revealing that a model may dominate specific task types even if its overall performance is poor. The authors apply their approach to the Visual Task Adaptation Benchmark (VTAB) to illustrate its practical usefulness.
The paper introduces a score‑calibrated robustness framework that transforms any fixed point predictor into a decision‑relevant uncertainty representation using distribution‑free conformal calibration. By employing the conformal score as the core unit of robustness, the authors derive both reliability‑based robust optimization and target‑oriented Conformal Robust Satisficing formulations, linking them through a shared robust decision frontier and a fragility measure. Experiments on synthetic data and a real online‑grocery inventory case study demonstrate the framework’s ability to improve reliability, reduce costs, and provide interpretable uncertainty scales for black‑box predictors.
arXiv:2606. 19569v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for reliable decision-making in safety-critical applications in probabilistic machine learning.
The paper introduces Calibration-Aware Uncertainty Cascades (CAUC), a post‑hoc framework that calibrates each model’s confidence independently and uses these calibrated scores to decide when to accept an early prediction, invoke a stronger model, or combine outputs. CAUC establishes a common reliability scale across heterogeneous models, decoupling deployment policies from specific model pools or budgets. Experiments on six language benchmarks show a 1.9% relative accuracy gain over strong‑model‑only inference while cutting strong‑model calls by about 47%, and on image classification it maintains or improves performance while reducing GFLOPs by up to 57%.
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.
arXiv:2603. 25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs).
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.
arXiv:2606. 30136v1 Announce Type: new Abstract: Humans facing algorithmic decision systems have been found to ``game'' them by altering their input data (at a cost to them) in order to favorably change the algorithmic outcomes they receive (at a cost to the algorithm).