arXiv:2604.12411v2 Announce Type: replace
Abstract: Segmentation models based on deep neural networks demonstrate strong generalization for medical image segmentation. However, they often exhibit ove...
By Qiuyu Tian, Haoliang Sun, Yunshan Wang, Yinghuan Shi, Yilong Yin
Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.
By Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Ma\v{s}ka, Ma\"elys Solal, Arya Yazdan-Panah, Vilma Bozgo, \"Omer S\"umer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim R\"adsch, Nadim Hammoud, Fruzsina Moln\'ar-G\'abor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jim\'enez-S\'anchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein
arXiv:2609.26384v1 Announce Type: new
Abstract: Medical image interpretation is high-volume and time-consuming, and while AI interpretation can reduce workload, fully autonomous deployment carries po...
By Emma Sun, Joshua Strong, Alison Noble
arXiv:2504. 19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance.
By Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts.
arXiv:2607. 25108v1 Announce Type: cross Abstract: Biomedical image analysis spans diverse modalities and tasks, yet real-world deployment is hindered by severe distribution shifts across scanners, protocols, and patient populations.
By Zihan Li, Feiyang Liu, Dandan Shan, Ruibo Wang, Qingqi Hong
arXiv:2602. 19502v2 Announce Type: replace Abstract: Agentic AI systems are increasingly capable of autonomous data science workflows, yet clinical prediction tasks demand domain expertise that purely automated approaches struggle to provide.
By Lalitha Pranathi Pulavarthy, Raajitha Muthyala, Aravind V Kuruvikkattil, Zhenan Yin, Rashmita Kudamala, Saptarshi Purkayastha
The study examines how three parameter‑efficient adaptation methods—linear heads on the raw CLS token, an MLP, and an attention‑pooling module—affect pathology classification accuracy and subgroup fairness when applied to a frozen Rad‑DINO chest X‑ray encoder. Using the MIMIC‑CXR dataset, the authors evaluate eight pathologies across race, sex, and imaging‑view subgroups, finding that attention pooling yields the best overall performance and encodes protected attributes most strongly, yet higher performance does not consistently reduce subgroup disparities. The results show that attribute encoding strength and layer choice do not reliably predict fairness outcomes, indicating that fairness must be assessed directly for each task.
By Dhruv Gupta, Emma A. M. Stanley, Fabio De Sousa Ribeiro, Sujal R. Desai, Ben Glocker
arXiv:2607. 10522v1 Announce Type: cross Abstract: Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback.
By Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
The paper introduces an agentic AI Scientist workflow that automates the entire baseline development process for medical imaging by combining literature-guided reasoning, automated code generation, and hypothesis-driven experimentation. Evaluated on four public benchmarks covering segmentation, classification, and detection, the pipeline consistently improves validation performance, achieving competitive leaderboard results such as 6th place on both PUMA tracks and 31st on MILK10k. The approach also shows strong domain generalization on MIDOG25 across scanners, tumor types, and species, demonstrating that a skill-based, literature-guided agentic workflow can reduce engineering effort without task-specific redesign.
By Eugenia Moris, Jos\'e Ignacio Orlando
arXiv:2607. 12464v1 Announce Type: cross Abstract: When labeled data are scarce, off-the-shelf diffusion models can augment training sets for few-shot medical image classification, but not all generated samples are equally useful for the downstream task.
By Jeeyung Kim, Erfan Esmaeili, Qiang Qiu
arXiv:2510. 07328v2 Announce Type: replace-cross Abstract: Medical decision systems increasingly rely on data from multiple sources to ensure reliable and unbiased diagnosis.
By Md Zubair, Hao Zheng, Grayson W. Armstrong, Lucy Q. Shen, Gabriela Wilson, Yu Tian, Xingquan Zhu