arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.
By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
arXiv:2608. 04766v1 Announce Type: cross Abstract: A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources.
By Bin Pu, Jiewen Yang, Liwen Wang, Ying Tan, Guannan He, Xingbo Dong, Qika Lin, Jiarong Guo, Lixian Yang, Zuozhu Liu, Shengli Li, Kenli Li
The paper introduces SonoCorpus, an open dataset of 456,963 ultrasound images with 1,626,085 expert masks from 53 public sources across 24 clinical applications and 17 countries, and SonoBase, an interactive segmentation foundation model pretrained on this data. SonoBase outperforms existing models (SAM2, MedSAM2, MedSAM3) on fifteen diverse evaluation datasets, matching specialist models and achieving clinically relevant accuracy for metrics such as ejection fraction, fetal head circumference, and gestational age. The authors provide full reproducibility resources, including checkpoints, optimizer states, and starter code, to enable community adoption and further development.
By Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan
arXiv:2606. 29586v1 Announce Type: cross Abstract: Vision-language foundation models have shown strong potential in medical image analysis.
By Hang Su, Chao Sun, Zhaofan Li, Wei Hu, Juhua Liu, Bo Du
arXiv:2608.21300v1 Announce Type: new
Abstract: Foundation models for medical image segmentation, like prompt-based MedSAM, generalize well across domains and modalities, often in zero or few-shot se...
By Marko Haralovi\'c, Sounic Akkaraju, Carlo Baretta, Vasil Zapryanov, Alexia Briassouli
arXiv:2606. 15176v1 Announce Type: cross Abstract: Ultrasound imaging is the most widely adopted medical modality globally due to its low cost and portability, yet artificial intelligence (AI) deployment remains constrained by reliance on GPU-accelerated models, creating a structural paradox where the cost of "intelligence" exceeds that of the imaging device itself.
By Weihao Gao
AnatoProto is a lightweight sequence‑level framework that adapts a frozen BiomedCLIP encoder for detecting the fetal abdominal circumference standard plane in low‑cost obstetric blind sweeps. It incorporates anatomy‑weighted spatial pooling, a within‑case prototype loss, a three‑stage cascade refinement, and a hybrid prediction head to address the highly imbalanced, short‑segment nature of the task. On the ACOUSLIC‑AI benchmark, AnatoProto achieves a test F1 of 67.72, surpassing the best foundation‑model baseline by 13.20 F1 and the best video temporal‑action‑detection baseline by 15.76 F1.
By Yuzhe Zhao
arXiv:2608. 14763v1 Announce Type: cross Abstract: Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement.
By Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi
The paper evaluates federated learning with Low‑Rank Adaptation (LoRA) for fine‑tuning the BiomedCLIP vision‑language model on chest X‑ray classification across four international cohorts. Federated LoRA improves shared‑class AUC from 0.687 to 0.802, outperforming isolated single‑cohort training and approaching a centralized reference. The study shows that SVD‑based product‑space aggregation (FlexLoRA) is crucial for performance, while FedProx offers no advantage over FedAvg in this setting.
By Sanjaya Poudel, Nirajan Kunwor, Manish Dhakal, Debesh Jha, Sunil Kumar Gaire
FreqDINO++ is a frequency‑guided multi‑task routing vision foundation model designed for universal ultrasound analysis. It introduces a Multi‑task Routing Adapter for efficient task‑common and task‑specific integration, a Frequency‑aware Feature Enhancer to capture multi‑scale frequency characteristics, and a Task‑aligned Collaborative Decoder that promotes collaboration between dense and global prediction tasks. Experiments on large‑scale multi‑task and external single‑task ultrasound benchmarks show that FreqDINO++ outperforms strong baselines and recent foundation models across 27 diverse clinical task scenarios, with promising generalization to unseen data.
By Qing Xu, Yixuan Zhang, Yue Li, Xiangjian He, Qian Zhang, Mainul Haque, Rong Qu, Wenting Duan, Jieyun Bai, Zhen Chen
arXiv:2607. 08219v2 Announce Type: replace-cross Abstract: The privacy requirements of medical data and its substantial variations across organs and modalities hinder the clinical implementation of medical AI.
By Junbin Mao, Xu Tian, Jianchun Zhu, Ludi Li, Jin Liu
This study introduces a pipeline that enhances ultrasound plane pose estimation for fetal brain imaging by providing continuous, real‑time proximity feedback to standard planes (SPs). It employs a semi‑supervised segmentation model achieving high mIoU scores on both SPs and non‑SPs, and integrates a classification step to filter out frames without the fetal brain. The system, validated on an NVIDIA Clara AGX edge device, runs at 39 Hz and has been tested on real scan videos from 17 sonographers, demonstrating its practical viability for clinical use.
By Chiara Di Vece, Antonio Cirigliano, Meala Le Lous, Raffaele Napolitano, Anna L. David, Donald Peebles, Pierre Jannin, Francisco Vasconcelos, Danail Stoyanov