Benchmarking competitions are central to AI development in medical imaging, but it is unclear if they provide representative, accessible, and reusable data for clinical relevance. This study systematically examined 249 challenges (458 tasks) across 19 imaging modalities, finding limited geographic, modality, and problem-type representation. Additionally, many datasets suffer from restrictive access, ambiguous licensing, and poor documentation, hindering reproducibility and long-term reuse.
Machine-generated by The Flow from the publisher's headline and feed description
— not written or checked by a human. The full article lives at arXiv Computer Vision.
The paper introduces an agentic AI Scientist workflow that automates the entire baseline development process for medical imaging by combining literature-guided reasoning, automated code generation, and hypothesis-driven experimentation. Evaluated on four public benchmarks covering segmentation, classification, and detection, the pipeline consistently improves validation performance, achieving competitive leaderboard results such as 6th place on both PUMA tracks and 31st on MILK10k. The approach also shows strong domain generalization on MIDOG25 across scanners, tumor types, and species, demonstrating that a skill-based, literature-guided agentic workflow can reduce engineering effort without task-specific redesign.
arXiv:2604. 26991v2 Announce Type: replace-cross Abstract: Machine learning models for medical image analysis often exhibit subgroup-dependent performance, which impacts how decisions should be allocated between automated systems and human experts under limited resources.
By Zheng Zhang, Milad Masroor, Cuong Nguyen, Tahir Hassan, Yuanhong Chen, David Rosewarne, Kevin Wells, Thanh-Toan Do, Gustavo Carneiro
arXiv:2504. 19621v2 Announce Type: replace Abstract: Machine learning (ML) systems for medical imaging have demonstrated remarkable diagnostic capabilities, but their susceptibility to biases poses significant risks, since biases may negatively impact generalization performance.
By Haroui Ma, Francesco Quinzan, Theresa Willem, Stefan Bauer
As vision-language models (VLMs) are increasingly applied to medical AI, existing benchmarks mainly focus on evaluating their diagnosis ability over given medical images and texts, implicitly assuming that standardized medical images, texts or question-answer pairs are already prepared. However, this assumption does not hold when we apply VLMs in real clinical practice, where medical data is often raw, heterogeneous, and fragmented across different sources.
arXiv:2607. 08219v2 Announce Type: replace-cross Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI.
By Junbin Mao, Xu Tian, Jianchun Zhu, Ludi Li, Jin Liu
arXiv:2607. 08219v1 Announce Type: cross Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI.
By Junbin Mao, Xu Tian, Jianchun Zhu, Ludi Li, Jin Liu