arXiv Machine Learning

Not Just How Much, But Where: Decomposing Epistemic Uncertainty into Per-Class Contributions

arXiv:2602. 21160v3 Announce Type: replace-cross Abstract: In safety-critical classification, the cost of failure is often asymmetric, yet Bayesian deep learning summarises epistemic uncertainty with a single scalar, mutual information (MI), that cannot distinguish whether a model's ignorance involves a benign or safety-critical class.

arXiv Machine Learning
2d ago

How Accurate Is Accurate Enough?

arXiv:2609.38785v1 Announce Type: new Abstract: How accurate must a numerical approximation be within a learning system? Primitive error alone cannot answer this question: errors of the same magnitud...

By Ningkang Peng, Qianfeng Yu, Jingyang Mao, Xiaoqian Peng, Yanhui Gu
arXiv Computer Vision
Aug 27

Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty

The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.

By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv AI
Sep 3

The Ceiling Is in the Channel: Auditing Learner Gaps and Measurement Frontiers in Clinical Prediction

The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.

By Sayeed Shafayet Chowdhury, Nusrat Jahan, Snehasis Mukhopadhyay, Shiaofen Fang, Vijay R. Ramakrishnan
arXiv Machine Learning
1d ago

A Ground-Truth Framework for Uncertainty Disentanglement with Posterior Risk

The paper introduces a ground‑truth framework for disentangling uncertainty into epistemic and aleatoric components using sample‑conditional pointwise posterior risk. It evaluates current methods, finding that Spectral‑normalized Neural Gaussian Processes and Variational Latent Gaussian Processes best recover the ground‑truth uncertainty, while most methods align more closely with posterior variance and miss predictor bias. The study also explores the entanglement of estimated uncertainties and the impact of modeling choices, providing practical guidance and releasing 13 semi‑synthetic datasets for further validation.

By Frieder Wizgall, Georg Tirpitz, Moritz Seiler, Kerstin Ritter, B\'alint Mucs\'anyi