arXiv Machine Learning

AI Alignment in Medical Imaging: Unveiling Hidden Biases Through Counterfactual Analysis

arXiv:2504. 19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance.

arXiv Computer Vision
Sep 3

Generating Medical Image Counterfactuals using Causal Explanations

The paper introduces a new framework for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from a classifier, the method deterministically produces edits within user-specified regions, requiring no additional training. Experiments on real-world medical imaging datasets show that these counterfactuals alter classifier predictions while staying closer to the original image than generative baselines, offering a clearer view of the model’s decision boundary.

By David A. Kelly, Tom Yaacov, Nathan Blake, Sander Beckers, Hana Chockler
Hugging Face Trending Papers
Sep 2

Generating Medical Image Counterfactuals using Causal Explanations

The paper introduces a new method for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from the classifier, the approach deterministically edits user-specified regions to alter predictions while staying closer to the original image than generative baselines. Experiments on real-world medical imaging datasets show that this technique provides a more direct and transparent view of the classifier’s decision boundary.

arXiv Machine Learning
Sep 11

Counterfactual Marginalisation: Framework for Evaluating Robustness to Nuisance Variables

The paper introduces counterfactual (CF) marginalisation, a test‑time evaluation method that assesses how robust classification models are to nuisance variables such as age or sex. By using a CF image generator to intervene on these parent variables, the method creates counterfactual versions of each test image and averages predictions over a chosen intervention distribution, yielding intervention‑aware predictions that filter out demographic effects while retaining patient‑specific latent information. These predictions are then used to define metrics for CF risk, calibration, stability, and worst‑case sensitivity, demonstrating the framework’s usefulness for quantitative robustness evaluation.

By Yasin Ibrahim, Hermione Warr, Robin J. Evans, Konstantinos Kamnitsas
arXiv Computer Vision
Aug 31

Medical Imaging AI Competitions Lack Fairness

Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.

By Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Ma\v{s}ka, Ma\"elys Solal, Arya Yazdan-Panah, Vilma Bozgo, \"Omer S\"umer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim R\"adsch, Nadim Hammoud, Fruzsina Moln\'ar-G\'abor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jim\'enez-S\'anchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein