arXiv AI

Measuring proximity to standard planes during fetal brain ultrasound scanning

This study introduces a pipeline that enhances ultrasound plane pose estimation for fetal brain imaging by providing continuous, real‑time proximity feedback to standard planes (SPs). It employs a semi‑supervised segmentation model achieving high mIoU scores on both SPs and non‑SPs, and integrates a classification step to filter out frames without the fetal brain. The system, validated on an NVIDIA Clara AGX edge device, runs at 39 Hz and has been tested on real scan videos from 17 sonographers, demonstrating its practical viability for clinical use.

arXiv Computer Vision
Aug 28

Anatomy-Guided Foundation Model Adaptation with Within-Case Prototype Supervision for Standard Plane Detection in Fetal Ultrasound Blind Sweeps

AnatoProto is a lightweight sequence‑level framework that adapts a frozen BiomedCLIP encoder for detecting the fetal abdominal circumference standard plane in low‑cost obstetric blind sweeps. It incorporates anatomy‑weighted spatial pooling, a within‑case prototype loss, a three‑stage cascade refinement, and a hybrid prediction head to address the highly imbalanced, short‑segment nature of the task. On the ACOUSLIC‑AI benchmark, AnatoProto achieves a test F1 of 67.72, surpassing the best foundation‑model baseline by 13.20 F1 and the best video temporal‑action‑detection baseline by 15.76 F1.

By Yuzhe Zhao
arXiv AI
Jul 22

FedCC: A Low-Resource Federated Adaptation of Foundation Models for Robust Corpus Callosum localization in Fetal Ultrasound Images

arXiv:2607. 18283v1 Announce Type: cross Abstract: Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnormalities.

By Alessandro Di Matteo, Sara Moccia, Giuseppe Rizzo, Gianpaolo Grisolia, Ricciarda Raffaelli, Lorenzo Vasciaveo, Francesco D'Antonio, Maria Chiara Fiorentino
arXiv Computer Vision
Aug 28

UniFLM: United Segmentation and Measurement on Fetal Limb Ultrasonic Image

The paper introduces UniFLM, a unified framework for segmenting and measuring fetal long bones in ultrasound images. It presents the Fetal Limb Bones (FLB) dataset with high‑quality annotations for the humerus, femur, tibia‑fibula, and radius‑ulna. UniFLM employs a Semantic‑Aware Skip Connection, a Positive Sampling strategy, and a Point Regression Mapping module to improve segmentation accuracy and bone length measurement, achieving superior performance over existing models on the FLB dataset.

By Zeen Zhou, Qiuhua Chen, Xiaojun Cao, Changmao Chen, Chao Sun, Bo Du
arXiv AI
Jun 10

FADA: Accessible fetal ultrasound interpretation and annotation with a selectively distilled unified vision-language model

arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.

By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
arXiv AI
Aug 18

Cross-Modal Ultrasound-MRI Learning for Fetal Brain Ventricular Volumetry and Abnormality Screening

arXiv:2608. 14763v1 Announce Type: cross Abstract: Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement.

By Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi
arXiv Computer Vision
Aug 25

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.

By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath