The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.
By Pranav Kaliaperumal, Manisha Kaliaperumal
arXiv:2610.01452v1 Announce Type: new
Abstract: While state-of-the-art automated models for medical image segmentation achieve high mean performance, they frequently suffer from localized, catastroph...
By Samuel Hart, Ahmad Yahya, Ahmed Karam Eldaly
arXiv:2607. 16317v1 Announce Type: cross Abstract: Deep networks now subtype brain tumors on MRI about as well as specialist readers, yet accuracy is not what keeps them out of the clinic.
By Medhansh Sharma
arXiv:2609.39429v1 Announce Type: cross
Abstract: Uncertainty Quantification (UQ) is a key requirement for trustworthy AI in high-stakes medical image analysis. In this work, we evaluate UQ in a mult...
By Gonzalo Esteban Mosquera Rojas, Sebastian R. van der Voort, Carolin M. Pirkl, Sandeep Kaushik, Marion Smits, Stefan Klein
The paper presents a pragmatic segmentation pipeline for brain metastases in the BraTS 2026 Task 1, using a 5‑fold nnU-Net ResEnc‑L ensemble trained for 1,000 epochs on 1,296 four‑modality cases. A rule‑based post‑processing cascade tuned for the lesion‑wise Dice similarity coefficient (LW‑DSC) improves performance, achieving LW‑DSC scores of 0.733, 0.751, 0.713, and 0.549 on enhancing tumour, tumour core, whole tumour, and resection cavity, respectively. The authors audit each post‑processing stage with a five‑fold out‑of‑fold analysis, confirm two stages as robust, and provide a mechanistic analysis of LW‑DSC, along with thirteen negative results that challenge common intuitions.
arXiv:2604. 15271v3 Announce Type: replace-cross Abstract: Reliable uncertainty estimation is critical for medical image segmentation, where automated contours feed downstream quantification and clinical decision support.
By Tianhao Fu, Austin Wang, Charles Chen, Roby Aldave-Garza, Yucheng Chen
The paper presents a segmentation pipeline for brain metastases in both pre‑ and post‑treatment cases using a 5‑fold nnU‑Net ResEnc‑L ensemble trained on 1,296 four‑modality cases. A rule‑based post‑processing cascade improves the lesion‑wise Dice similarity coefficient (LW‑DSC) for enhancing tumour, tumour core, whole tumour, and resection cavity sub‑regions, achieving LW‑DSC scores of 0.733, 0.751, 0.713, and 0.549 respectively on the official validation leaderboard. The authors conduct a five‑fold out‑of‑fold analysis to validate the robustness of each post‑processing stage, provide a mechanistic explanation of LW‑DSC behaviour, and report thirteen negative results that challenge common intuitions, with all code released under Apache‑2.0.
By Haobin Liu, Xin Wang
The paper introduces SWIFT, a Swin V2‑based model pretrained on 10,444 3D CT volumes and fine‑tuned for rectal cancer segmentation on T2‑weighted MRI. Four configurations—full fine‑tuning (SWIFT), decoder compression (SWIFTe), low‑rank adaptation (SWIFTe‑LoRA), and a LoRA‑decoder ensemble (SWIFTe‑LDE4)—were evaluated on 247 cases, showing that SWIFTe reduces parameters by 70.1% while improving tumor detection and radiomic agreement. The study also demonstrates a trade‑off between detection and boundary agreement, and highlights that SWIFTe‑LDE4 achieves the lowest calibration errors after temperature scaling.
By Aneesh Rangnekar, Jorge Tapias Gomez, Joseph O Deasy, Harini Veeraraghavan
arXiv:2609.24769v1 Announce Type: new
Abstract: Brain metastases are the most common intracranial malignancy, occurring in roughly 30% of patients with primary solid tumors and carrying a median surv...
By Mahdi Islam, Musarrat Tabassum
arXiv:2608.28681v1 Announce Type: new
Abstract: Probability calibration aligns model confidence with predictive accuracy, enabling clinicians to identify unreliable segmentation regions. This alignme...
By Jiaheng Dai, Weidong Guo, Qingbiao Li, Jie Xu, Yi Guo, Yuanyuan Wang, Zeju Li
arXiv:2608. 14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity.
By Simon Baur, Arne Schernich, Ekin B\"oke, Wojciech Samek, Jackie Ma