Evaluating Epistemic Uncertainty: Beyond OOD Detection and Active Learning
arXiv:2607. 14817v1 Announce Type: cross Abstract: Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning.
arXiv:2606. 26990v1 Announce Type: cross Abstract: Uncertainty estimates in machine learning are typically evaluated using generic metrics such as the negative log-likelihood and expected calibration error, yet good performance on such metrics does not necessarily imply high utility in downstream decisions.
arXiv:2607. 14817v1 Announce Type: cross Abstract: Current evaluation of epistemic uncertainty relies on tasks such as out-ofdistribution detection and active learning.
arXiv:2504. 18433v3 Announce Type: replace Abstract: Uncertainty quantification is crucial in machine learning, yet most (axiomatic) studies of uncertainty measures focus on classification, leaving a gap in regression settings with limited formal justification and evaluations.
arXiv:2606. 19569v1 Announce Type: new Abstract: Uncertainty quantification (UQ) is essential for reliable decision-making in safety-critical applications in probabilistic machine learning.
arXiv:2606. 10777v1 Announce Type: new Abstract: Uncertainty estimation is critical for deploying machine learning models in high-stakes settings.
arXiv:2607. 27143v1 Announce Type: new Abstract: High-stakes decision systems in credit scoring, fraud detection, healthcare, and industrial safety require reliable uncertainty quantification under severe class imbalance and asymmetric error costs.
arXiv:2603. 25670v3 Announce Type: replace Abstract: Safety monitoring is essential for Cyber-Physical Systems (CPSs).
arXiv:2606. 05551v1 Announce Type: cross Abstract: Reliable decision making pipelines powered by machine learning models require uncertainty quantification (UQ) methods that come with explicit safety guarantees.
arXiv:2607. 16239v1 Announce Type: new Abstract: AI judges offer a scalable, low-cost alternative to human evaluation, but their outputs can be biased relative to human preferences and highly item-dependent, varying across judges, tasks, and domains.
arXiv:2606. 30136v1 Announce Type: new Abstract: Humans facing algorithmic decision systems have been found to ``game'' them by altering their input data (at a cost to them) in order to favorably change the algorithmic outcomes they receive (at a cost to the algorithm).
Probabilistic models are typically trained using task-agnostic objectives like log-loss, which can lead to significant errors in downstream estimation. This disconnect is especially critical in Inverse Probability Weighting (IPW) for causal inference, where propensity score errors near $0$ and $1$ often lead to high bias and variance.
arXiv:2510. 25599v2 Announce Type: replace Abstract: Regression tasks, notably in safety-critical domains, require reliable uncertainty quantification, yet the literature remains largely classification-focused.
Using LLMs as judges has become standard practice for evaluating model outputs at scale. This is particularly common for subjective, open-ended tasks such as assessing helpfulness or alignment, where no single reference answer exists.