The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.
By Pranav Kaliaperumal, Manisha Kaliaperumal
arXiv:2610.01452v1 Announce Type: new
Abstract: While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastroph...
By Samuel Hart, Ahmad Yahya, Ahmed Karam Eldaly
arXiv:2607. 16317v1 Announce Type: cross Abstract: Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic.
By Medhansh Sharma
arXiv:2609.39429v1 Announce Type: cross
Abstract: Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a mult...
By Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
The paper presents a pragmatic segmentation pipeline for brain metastases in the BraTS 2026 Task 1, using a 5‑fold nnU-Net ResEnc‑L ensemble trained for 1,000 epochs on 1,296 four‑modality cases. A rule‑based post‑processing cascade tuned for the lesion‑wise Dice similarity coefficient (LW‑DSC) improves performance, achieving LW‑DSC scores of 0.733, 0.751, 0.713, and 0.549 on enhancing tumour, tumour core, whole tumour, and resection cavity, respectively. The authors audit each post‑processing stage with a five‑fold out‑of‑fold analysis, confirm two stages as robust, and provide a mechanistic analysis of LW‑DSC, along with thirteen negative results that challenge common intuitions.