arXiv AI

Towards Trustworthy AI for Glioma Diagnosis: A Task-Aware Evaluation of Uncertainty Quantification

arXiv Machine Learning
Jul 28

Trustworthy Medical Segmentation: Uncertainty-Aware U-Net Evaluation Under Clinical Image Degradation

arXiv:2607. 22727v1 Announce Type: cross Abstract: Medical image segmentation models often report high benchmark accuracy under ideal imaging conditions, yet their failures under clinical degradation can be quiet: sensor noise, patient motion, low- resolution acquisition, and contrast variability may all alter model behavior without producing an obvious warning.

By Pranav Kaliaperumal, Manisha Kaliaperumal
arXiv Machine Learning
Aug 18

Beyond Boundary Noise: Aggregated Aleatoric Uncertainty Fails to Capture Presence Ambiguity in 3D Lung Nodule Segmentation

arXiv:2608. 14766v1 Announce Type: cross Abstract: Uncertainty estimation is critical for the safe clinical deployment of deep learning in medical image segmentation, with aleatoric uncertainty theoretically designed to capture irreducible data ambiguity.

By Simon Baur, Arne Schernich, Ekin B\"oke, Wojciech Samek, Jackie Ma
arXiv Computer Vision
Aug 27

Reliability analysis for BraTS-GoAT segmentation: a controlled robustness study of deep-ensemble uncertainty

The study evaluates the reliability of deep‑ensemble uncertainty for brain tumour segmentation on the BraTS‑GoAT dataset. A 5‑fold cross‑validated nnU‑Net baseline and a 3‑seed deep ensemble were compared for calibration and error detection; the ensemble showed modest gains in calibration on in‑distribution data but the single model’s confidence remained flat while accuracy degraded under synthetic corruptions. Disagreement among ensemble members rose sharply with corruption severity, proving to be a more sensitive indicator of acquisition shift than single‑model confidence.

By Riya Deepak Shet, Chenxi Liang, Le Zhang
arXiv AI
Sep 16

A Vision-Language Foundation Model for Precise and Comprehensive Brain Tumor Diagnosis from Preoperative Multimodal Data

arXiv:2609.16597v1 Announce Type: cross Abstract: Background Non-invasive presurgical diagnosis of brain tumor types from Magnetic Resonance Imaging (MRI) is essential but challenging due to overlapp...

By Yinong Wang (Joyce), Jianwen Chen (Joyce), Zhou Chen (Joyce), Shuwen Kuang (Joyce), Haoning Jiang (Joyce), Yanzhao Shi (Joyce), Huichun Yuan (Joyce), Yan-ran (Joyce), Wang, Bing Wang, Lei Wu, Bin Tang, Li Meng, Baihua Luo, Bin Zhou, Wei Ding, Weiming Zhong, Wei Hou, Yuanbing Chen, Zhiping Wan, Wei Wang, Zhenkun Xiao, Wenwu Wan, Allen He, Yuyin Zhou, Longbo Zhang, Feifei Wang, Zhixiong Liu, Michael Iv, Xuan Gong, Liangqiong Qu
arXiv Machine Learning
Aug 27

Performance uncertainty in medical image analysis: a large-scale investigation of confidence intervals

The study investigates how confidence intervals (CIs) behave in medical imaging AI by analyzing 24 segmentation and classification tasks with 19 models per task, various metrics, aggregation strategies, and CI methods. It finds that required sample sizes for reliable CIs vary widely, CI behavior depends on performance metrics, aggregation strategy, and problem type, and that different CI methods differ in reliability and precision. The authors provide a decision tree to guide researchers in selecting appropriate CI methods, aiming to support future consensus guidelines on reporting performance uncertainty.

By Pascaline Andr\'e (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Charles Heitz (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Evangelia Christodoulou (German Cancer Research Center), Annika Reinke (German Cancer Research Center), Carole H. Sudre (Unit for Lifelong Health and Ageing at UCL, Department of Population Science and Experimental Medicine and Hawkes InstituteCentre for Medical Image Computing, Department of Computer Science, University College London, UK), Michela Antonelli (School of Biomedical Engineering and Imaging Science, King's College London, UK), Patrick Godau (German Cancer Research Center), M. Jorge Cardoso (School of Biomedical Engineering and Imaging Science, King's College London, UK), Antoine Gilson (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Sophie Tezenas du Montcel (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France), Ga\"el Varoquaux (SODA project team, Inria Saclay-\^Ile-de-France, France), Lena Maier-Hein (German Cancer Research Center), Olivier Colliot (Sorbonne Universit\'e, Institut du Cerveau - Paris Brain Institute - ICM, CNRS, Inria, Inserm, AP-HP, H\^opital de la Piti\'e-Salp\^etri\`ere, Paris, France)
arXiv Machine Learning
Jun 30

TRACE: A Concept Bottleneck Model for Longitudinal 3D Glioblastoma Response Assessment

arXiv:2606. 30313v1 Announce Type: cross Abstract: Longitudinal glioblastoma response assessment requires comparing subtle tumor changes across MRI time points using structured clinical criteria such as RANO.

By Alia Tarek, Hamsa Saberr, Hamza Elghonemy, Youssef Afify, Tamer Basha, Omair Shahzad Bhatti, Abdulrahman M. Selim, Hasan Md Tusfiqur Alam Daniel Sonntag
arXiv Computer Vision
Sep 21

Uncertainty-driven training for three-dimensional calibrated lung nodule classification

The paper introduces an uncertainty‑driven training framework for 3D CT lung nodule classification that uses validation‑based uncertainty estimates to reweight the loss, aiming to improve predictive performance and probability calibration. Two uncertainty quantification methods—Monte Carlo Dropout and Evidential Deep Learning—are evaluated across multiple backbone architectures (ResNet, DenseNet, EfficientNet, ViT, Swin) on the LIDC‑IDRI and NoduleMNIST3D datasets. The approach yields comparable classification accuracy to conventional training while substantially reducing expected calibration error, especially on convolutional backbones, and shows that simple temperature scaling can also achieve strong calibration.

By Giuseppe Tripodi, Alessandro De Rosis, Saleh Rezaeiravesh