The paper introduces an agentic AI Scientist workflow that automates the entire baseline development process for medical imaging by combining literature-guided reasoning, automated code generation, and hypothesis-driven experimentation. Evaluated on four public benchmarks covering segmentation, classification, and detection, the pipeline consistently improves validation performance, achieving competitive leaderboard results such as 6th place on both PUMA tracks and 31st on MILK10k. The approach also shows strong domain generalization on MIDOG25 across scanners, tumor types, and species, demonstrating that a skill-based, literature-guided agentic workflow can reduce engineering effort without task-specific redesign.
By Eugenia Moris, Jos\'e Ignacio Orlando
Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.
By Annika Reinke, Evangelia Christodoulou, Sthuthi Sadananda, A. Emre Kavur, Khrystyna Faryna, Daan Schouten, Bennett A. Landman, Carole Sudre, Olivier Colliot, Nick Heller, Sophie Loizillon, Martin Ma\v{s}ka, Ma\"elys Solal, Arya Yazdan-Panah, Vilma Bozgo, \"Omer S\"umer, Siem de Jong, Sophie Fischer, Michal Kozubek, Tim R\"adsch, Nadim Hammoud, Fruzsina Moln\'ar-G\'abor, Steven Hicks, Michael A. Riegler, Anindo Saha, Vajira Thambawita, Pal Halvorsen, Amelia Jim\'enez-S\'anchez, Qingyang Yang, Veronika Cheplygina, Sabrina Bottazzi, Alexander Seitel, Spyridon Bakas, Alexandros Karargyris, Kiran Vaidhya Venkadesh, Bram van Ginneken, Lena Maier-Hein
Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback. Translating this capability to medical imaging remains difficult because each task imposes modality-specific experimentation and strict requirements for validation protocols and prediction artifacts.
arXiv:2607. 10522v1 Announce Type: cross Abstract: Large language model (LLM) agents are beginning to automate machine learning engineering (MLE) by coupling planning, code execution, debugging, and empirical feedback.
By Shengyuan Liu, Jia-Xuan Jiang, Boyun Zheng, Cheng Wang, Zipei Wang, Wentao Pan, Hongtao Wu, Houwen Peng, Yu Gu, Lichao Sun, Yixuan Yuan
The paper critiques the prevailing model-first approach in AI-driven image processing, arguing that researchers often prioritize benchmark performance over genuine understanding of real-world imaging problems. It proposes a problem-first framework that separates the physical imaging issue, solution principle, statistical estimator, and computational implementation, and introduces a six-stage workflow to guide research from problem formulation to evaluation. Case studies in super-resolution and low-light enhancement illustrate how benchmark datasets can misrepresent real tasks and emphasize the need for clearer standards on evidence, reproducibility, and uncertainty.
By Guoping Qiu
arXiv:2406. 11868v2 Announce Type: replace-cross Abstract: The emergence of foundational models represents a paradigm shift in medical imaging, offering extraordinary capabilities in disease detection, diagnosis, and treatment planning.
By Debesh Jha, Gorkem Durak, Abhijit Das, Jasmer Sanjotra, Onkar Susladkar, Suramyaa Sarkar, Ashish Rauniyar, Nikhil Kumar Tomar, Linkai Peng, Sirui Li, Koushik Biswas, Ertugrul Aktas, Elif Keles, Matthew Antalek, Zheyuan Zhang, Bin Wang, Xin Zhu, Hongyi Pan, Deniz Seyithanoglu, Alpay Medetalibeyoglu, Vanshali Sharma, Vedat Cicek, Amir A. Rahsepar, Rutger Hendrix, A. Enis Cetin, Bulent Aydogan, Mohamed Abazeed, Frank H. Miller, Rajesh N. Keswani, Hatice Savas, Sachin Jambawalikar, Daniela P. Ladner, Amir A. Borhani, Concetto Spampinato, Michael B. Wallace, Ulas Bagci
The article reports on the deployment of imaging AI across six hospitals using the open, self‑hosted PACS‑AI platform. It emphasizes that the main limitation is not model accuracy but the infrastructure needed to route studies, display results, collect feedback, and audit runs. In one center, angiography models processed 84.8% of jobs, with failures mainly due to missing diagnostic views, and 78.1% of clinician ratings were positive.
By Samuel Kadoury, Julie G. Hussin, Pascal Th\'eriault-Lauzier, Laurent L\'etourneau-Guillon, Rob Lewis, Adam McArthur, Gordon J. Harris, Houda Bahig, Pierre-Luc D\'eziel, Jay Kshirsagar, Jacob L. Jaremko, Julien Cohen-Adad, Jacques Delfrate, Robert Avram
arXiv:2607. 03634v1 Announce Type: new Abstract: Artificial intelligence (AI) has achieved extraordinary capabilities despite lacking many of the conceptual and scientific foundations associated with mature disciplines.
By Timothy Nguyen
Methodological progress in artificial intelligence (AI) for electronic health records (EHRs) depends on our ability to determine which algorithms work better, and under which conditions. However, such...
arXiv:2603. 27341v4 Announce Type: replace Abstract: Recent Artificial Intelligence (AI) models have matched or exceeded human experts in several benchmarks of biomedical task performance, but surgical benchmarks in particular are often missing from prominent medical benchmark suites.
By Kirill Skobelev, Eric Fithian, Yegor Baranovski, Jack Cook, Sandeep Angara, Shauna Otto, Zhuang-Fang Yi, John Zhu, Neeraj Mainkar, Margaux Masson-Forsythe, Daniel A. Donoho, X. Y. Han
The study re‑implements 12 AI algorithms for electronic health records within a unified framework and evaluates them on MIMIC‑IV and NWICU datasets. It compares expert‑authored clinically meaningful tasks with randomly generated tasks, finding that pairwise algorithm comparisons transfer well across task families and datasets, yet clinically meaningful tasks show stronger task‑method interactions. The results also reveal that newer algorithms do not consistently outperform older ones, with gradient‑boosted trees remaining highly competitive when combined with modern EHR representations.
By Florent Pollet, Matthew McDermott
arXiv:2608.28820v1 Announce Type: new
Abstract: Computational pathology (CompPath) is transforming medicine by leveraging artificial intelligence (AI) algorithms to support diagnosis, prognosis, and...
By Shubham Innani, Suhang You, Adam Shephard, Bhakti Baheti, Francesco Ciompi, Joe Yeong, Nasir Rajpoot, Michael Feldman, Solene Florence Kammerer-Jacquet, Dimitrios Makris, Geert Litjens, Anne L. Martel, Jana Lipkova, April Khademi, Spyridon Bakas, for the MICCAI SIG-CompPath