Confidence is Not Reliability: Rethinking MC Dropout in Brain Tumour Segmentation
arXiv:2606. 19300v1 Announce Type: cross Abstract: Glioma segmentation in multiparametric MRI is a critical component of treatment planning.
The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
arXiv:2606. 19300v1 Announce Type: cross Abstract: Glioma segmentation in multiparametric MRI is a critical component of treatment planning.
arXiv:2607. 16317v1 Announce Type: cross Abstract: Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic.
arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.
arXiv:2608. 09768v1 Announce Type: new Abstract: A prediction that is both confident and wrong is a critical reliability failure because it can bypass abstention and human review precisely when the model is mistaken.
arXiv:2608. 14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity.
The study investigates how confidence intervals (CIs) behave in medical imaging AI by analyzing 24 segmentation and classification tasks with 19 models per task, various metrics, aggregation strategies, and CI methods. It finds that required sample sizes for reliable CIs vary widely, CI behavior depends on performance metrics, aggregation strategy, and problem type, and that different CI methods differ in reliability and precision. The authors provide a decision tree to guide researchers in selecting appropriate CI methods, aiming to support future consensus guidelines on reporting performance uncertainty.
arXiv:2608.28681v1 Announce Type: new Abstract: Probability calibration aligns model confidence with predictive accuracy, enabling clinicians to identify unreliable segmentation regions. This alignme...
arXiv:2604. 15271v3 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decision support.
arXiv:2609.02600v1 Announce Type: new Abstract: This work presents an approach to the Generalizability Across Tumors (BraTS-GoAT) task of the BraTS 2026 Challenge, which focuses on robust segmentatio...
arXiv:2608. 14768v1 Announce Type: cross Abstract: Skin lesion classifiers can be confidently wrong on the cases that matter most, so knowing when a prediction should not be trusted is clinically as useful as the prediction.
arXiv:2509. 05238v2 Announce Type: replace-cross Abstract: Deep learning (DL) has transformed neuroimaging by delivering state-of-the-art performance with reduced computation times.
The paper introduces a framework that distinguishes two causes of saturation in clinical prediction: a learner gap, where the model fails to use available information, and a measurement‑channel ceiling, where the recorded variables limit performance. It provides theoretical characterizations, finite‑sample diagnostics, and empirical audits across three large cohorts, showing that well‑tuned models approach the frontier while deficient learners leave large gaps. A PRISMA‑guided synthesis across 104 tasks reveals consistent channel‑level patterns, suggesting that improving the learner or the measurement channel can audit and potentially lift performance.