arXiv Machine Learning By Nathan Le, Magdalini Paschali, Arogya Koirala, Andrew Johnston, Zhongnan Fang, David B. Larson, Akshay S. Chaudhari, Camila Gonzalez

Improving Calibration of Black-Box Radiology AI Using Test-Time Augmentation

Read the original on arXiv Machine Learning →

The paper presents DualTTA, a model‑agnostic framework that improves the calibration of black‑box radiology AI systems by applying clinically grounded test‑time augmentations (geometric and physics‑inspired 3D CT perturbations) and learning probability‑level aggregation strategies. Without accessing model internals or training data, DualTTA achieved the best overall calibration across pulmonary embolism and intracranial hemorrhage detection tasks, reducing Expected Calibration Error by 54% and 43% respectively. It also outperformed traditional uncertainty estimation methods that require internal model access, such as Temperature Scaling, MC Dropout, and Deep Ensembles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Sep 21

Uncertainty-driven training for three-dimensional calibrated lung nodule classification

The paper introduces an uncertainty‑driven training framework for 3D CT lung nodule classification that uses validation‑based uncertainty estimates to reweight the loss, aiming to improve predictive performance and probability calibration. Two uncertainty quantification methods—Monte Carlo Dropout and Evidential Deep Learning—are evaluated across multiple backbone architectures (ResNet, DenseNet, EfficientNet, ViT, Swin) on the LIDC‑IDRI and NoduleMNIST3D datasets. The approach yields comparable classification accuracy to conventional training while substantially reducing expected calibration error, especially on convolutional backbones, and shows that simple temperature scaling can also achieve strong calibration.

By Giuseppe Tripodi, Alessandro De Rosis, Saleh Rezaeiravesh
arXiv Machine Learning
Jun 16

A Multi-Center Benchmark for Abdominal Disease Diagnosis and Report Generation from Non-Contrast CT

arXiv:2606. 16991v1 Announce Type: cross Abstract: Multiphasic contrast-enhanced CT (CECT) is widely used for abdominal lesion characterization, yet it carries inherent risks of contrast-induced nephropathy, escalates acquisition burden, and heavily contributes to radiologist workload.

By Mariam Elbakry, Aliaa Sayed Sheha, Salma Hassan Tantawy, Aya Yassin, Concetto Spampinato, Karim Lekadir, Xiaomeng Li, Marawan Elbatel
arXiv AI
Jul 8

CONFLUX: A Latent Diffusion Model for 3D Chest-CT Synthesis with RL Post-Training

arXiv:2607. 02998v2 Announce Type: replace-cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.

By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv AI
Jul 7

CONFLUX: A Latent Diusion Model for 3D Chest-CT Synthesis with RL Post-Training

arXiv:2607. 02998v1 Announce Type: cross Abstract: Controllable generative models of 3D medical images can synthesize volumes with specified clinical attributes, but this demands samples that are simultaneously high-fidelity, natively 3D, and faithful to the requested conditioning.

By Max Van Puyvelde, Halil Ibrahim Gulluk, Wim Van Criekinge, Olivier Gevaert
arXiv Computer Vision
Sep 24

nnFoundation: 3D Foundation Models for Radiology

nnFoundation introduces complementary convolutional and transformer-based 3D foundation models for radiology, trained on 2.1 million CT, MRI, and PET volumes from 125 datasets. The models are evaluated on 108 tasks—including segmentation, detection, classification, report generation, and image retrieval—under domain shift, low-data, and low-compute scenarios, consistently outperforming prior 3D foundation models and training from scratch. Performance varies by task type, with convolutional models excelling at spatially localized tasks and transformer models at global semantic reasoning, and dynamic alignment with dataset characteristics further enhances transferability.

By Constantin Ulrich Harsy, Tassilo Wald, Karol Gotkowski, Yannick Kirchhoff, Marcel Knopp, Maximilian Rokuss, Elisa Stegmeier, Philipp Schader, Dasha Trofimova, Raphael Stock, Kim-Celine Kahl, Stephen Schaumann, Selen Erkan, David Zimmerer, Stefan Denner, Moritz Langenberg, Sebastian Ziegler, Katharina Eckstein, Maximilian Fischer, Jonathan Suprijadi, B\'alint Kov\'acs, Benjamin Hamm, Anand Deshpande, Dimitrios Bounias, Nico Disch, Shuhan Xiao, Jessica K\"achele, Jan Sellner, Rajesh Baidya, Jeremias Traub, Lars Kr\"amer, Maximilian Zenk, Tim R\"adsch, Stefan Dvoretskii, Robin Peretzke, Jonathan Deissler, Alexandra Ertl, Partha Ghosh, Kris Dreher, Stefan Dinkelacker, Annika Reinke, Evangelia Christodoulou, Numan Saeed, Yoland Savriama, Santiago Estrada, David K\"ugler, Laura Alexandra Daza Barragan, Cristina Isabel Gonzalez Osorio, Jan Peeken, Michael Baumgartner, Marvin Teichmann, Guillaume Chabin, Matthias Kirchler, Valentin Koch, for the ALFA study, Markus Hohenhaus, Dimitri Koslov, Nina Decker, Mohammad Yaqub, Arnd Heuser, Martin Reuter, Julia A. Schnabel, Tobias Heimann, Florin Ghesu, Paul Brachmann, Claus P. Heu{\ss}el, Alexander Radbruch, Gianluca Brugnara, Aditya Rastogi, Martha Foltyn-Dumitru, Heinz-Peter Schlemmer, Ignaz Reicht, Julius C. Holzschuh, Michael Bach, Bram Stieltjes, Kai Schlamp, Lena Maier-Hein, Marco Nolden, Ralf Floca, Paul F. J\"ager, Philipp Vollmuth, Fabian Isensee, Klaus H. Maier-Hein