arXiv:2606. 15837v1 Announce Type: cross Abstract: Deep neural networks (DNNs) frequently fail to generalize to out-of-distribution (OOD) medical images because of variations in scanners and acquisition protocols.
By Jimut B. Pal, Suyash P. Awate
arXiv:2609.06729v2 Announce Type: replace
Abstract: Accurate 3D brain tumour segmentation from multi-modal Magnetic Resonance Imaging (MRI) is essential for clinical diagnosis and treatment planning....
By Libing Kuang, Soren Salehi, Ziling Wu, Ahmad P. Tafti, Armaghan Moemeni
arXiv:2608. 14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity.
By Simon Baur, Arne Schernich, Ekin B\"oke, Wojciech Samek, Jackie Ma
The paper introduces an uncertainty‑driven training framework for 3D CT lung nodule classification that uses validation‑based uncertainty estimates to reweight the loss, aiming to improve predictive performance and probability calibration. Two uncertainty quantification methods—Monte Carlo Dropout and Evidential Deep Learning—are evaluated across multiple backbone architectures (ResNet, DenseNet, EfficientNet, ViT, Swin) on the LIDC‑IDRI and NoduleMNIST3D datasets. The approach yields comparable classification accuracy to conventional training while substantially reducing expected calibration error, especially on convolutional backbones, and shows that simple temperature scaling can also achieve strong calibration.
By Giuseppe Tripodi, Alessandro De Rosis, Saleh Rezaeiravesh
arXiv:2609.06807v1 Announce Type: cross
Abstract: In this work, we comprehensively evaluate three popular feature-extraction paradigms in AI-based neuroimaging modeling: (1) computation of anatomical...
By Boyang Yu, Miquel Lopez Escoriza, Long Chen, Arjun V. Masurkar, Narges Razavian, Carlos Fernandez-Granda
The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.
By Riya Deepak Shet, Chenxi Liang, Le Zhang