The paper introduces a new framework for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from a classifier, the method deterministically produces edits within user-specified regions, requiring no additional training. Experiments on real-world medical imaging datasets show that these counterfactuals alter classifier predictions while staying closer to the original image than generative baselines, offering a clearer view of the model’s decision boundary.
By David A. Kelly, Tom Yaacov, Nathan Blake, Sander Beckers, Hana Chockler
arXiv:2605.10894v2 Announce Type: replace
Abstract: Deep learning models in medical imaging often fail when deployed in new clinical environments due to distribution shifts in demographics, scanner h...
By Moritz Stammel, Fabio De Sousa Ribeiro, Raghav Mehta, M\'elanie Roschewitz, Ben Glocker
The paper introduces a new method for generating medical image counterfactuals that does not rely on auxiliary generative models. By extracting causal evidence directly from the classifier, the approach deterministically edits user-specified regions to alter predictions while staying closer to the original image than generative baselines. Experiments on real-world medical imaging datasets show that this technique provides a more direct and transparent view of the classifier’s decision boundary.
arXiv:2609.24879v1 Announce Type: new
Abstract: Counterfactual image generation answers questions about how a subject would have looked under retrospective, hypothetical scenarios. Recent methods hav...
By Xiaodan Xing, Rajat R. Rasal, Julia A. Meister, Sara Ghorayeb, Galvin Khara, Jessica Schrouff
arXiv:2609.26623v1 Announce Type: new
Abstract: Diffusion-based synthetic data generation offers a promising route for sharing medical imaging data without releasing sensitive patient records. Howeve...
By Mischa Dombrowski, Bernhard Kainz
The paper introduces counterfactual (CF) marginalisation, a test‑time evaluation method that assesses how robust classification models are to nuisance variables such as age or sex. By using a CF image generator to intervene on these parent variables, the method creates counterfactual versions of each test image and averages predictions over a chosen intervention distribution, yielding intervention‑aware predictions that filter out demographic effects while retaining patient‑specific latent information. These predictions are then used to define metrics for CF risk, calibration, stability, and worst‑case sensitivity, demonstrating the framework’s usefulness for quantitative robustness evaluation.
By Yasin Ibrahim, Hermione Warr, Robin J. Evans, Konstantinos Kamnitsas
arXiv:2506.10633v2 Announce Type: replace
Abstract: Latent Diffusion Models have shown remarkable results in text-guided image synthesis in recent years. In the domain of natural (RGB) images, recent...
By Konstantinos Vilouras, Ilias Stogiannidis, Junyu Yan, Alison Q. O'Neil, Sotirios A. Tsaftaris
Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.
By Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Ma\v{s}ka, Ma\"elys Solal, Arya Yazdan-Panah, Vilma Bozgo, \"Omer S\"umer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim R\"adsch, Nadim Hammoud, Fruzsina Moln\'ar-G\'abor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jim\'enez-S\'anchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein
arXiv:2608.29456v1 Announce Type: new
Abstract: As artificial intelligence is increasingly integrated into chest X-ray (CXR) interpretation, triage, and clinical decision support, understanding its v...
By Basudha Pal, Arjun Narayanan, Neha Ajith, Vikas R Bhat, Muhammad Umair
arXiv:2603.08459v2 Announce Type: replace
Abstract: Safe predictions are a crucial requirement for integrating predictive models into clinical decision support systems. One approach to improving trus...
By L. Juli\'an Lechuga L\'opez, Tim G. J. Rudner, Farah E. Shamout
arXiv:2607. 02596v1 Announce Type: cross Abstract: Deep learning models for medical diagnosis frequently exhibit substantial performance disparities across sensitive subgroups (e.
By Xinyu Jia, Weidong Guo, Wangyuan Zhao, Yi Guo, Zeju Li, Yuanyuan Wang
arXiv:2512. 09185v4 Announce Type: replace-cross Abstract: Understanding disease progression is a central clinical challenge with direct implications for early diagnosis and personalized treatment.
By Hao Chen, Rui Yin, Yifan Chen, Qi Chen, Chao Li