AnatoProto is a lightweight sequence‑level framework that adapts a frozen BiomedCLIP encoder for detecting the fetal abdominal circumference standard plane in low‑cost obstetric blind sweeps. It incorporates anatomy‑weighted spatial pooling, a within‑case prototype loss, a three‑stage cascade refinement, and a hybrid prediction head to address the highly imbalanced, short‑segment nature of the task. On the ACOUSLIC‑AI benchmark, AnatoProto achieves a test F1 of 67.72, surpassing the best foundation‑model baseline by 13.20 F1 and the best video temporal‑action‑detection baseline by 15.76 F1.
By Yuzhe Zhao
arXiv:2608. 07116v1 Announce Type: cross Abstract: Camera localization in bronchoscopy remains a challenging problem due to stringent accuracy requirements, real-time constraints, and limited training data.
By Lumin Chen, Qingyao Tian, Jinpeng Li, Haoyu Jiang, Huai Liao, Xinyan Huang, Hongbin Liu, Dong Yi
arXiv:2607. 18283v1 Announce Type: cross Abstract: Accurate localization of the corpus callosum (CC) in fetal ultrasound (US) images is crucial for the early identification of neurodevelopmental abnormalities.
By Alessandro Di Matteo, Sara Moccia, Giuseppe Rizzo, Gianpaolo Grisolia, Ricciarda Raffaelli, Lorenzo Vasciaveo, Francesco D'Antonio, Maria Chiara Fiorentino
arXiv:2608.28715v1 Announce Type: cross
Abstract: Ultrasound-guided interventions can require localization of an untracked 2D frame within a 3D anatomical reference. Rigid slice-to-volume registratio...
By Niklas Schwarz, Jens Kleesiek, Moritz Rempe
The paper introduces UniFLM, a unified framework for segmenting and measuring fetal long bones in ultrasound images. It presents the Fetal Limb Bones (FLB) dataset with high‑quality annotations for the humerus, femur, tibia‑fibula, and radius‑ulna. UniFLM employs a Semantic‑Aware Skip Connection, a Positive Sampling strategy, and a Point Regression Mapping module to improve segmentation accuracy and bone length measurement, achieving superior performance over existing models on the FLB dataset.
By Zeen Zhou, Qiuhua Chen, Xiaojun Cao, Changmao Chen, Chao Sun, Bo Du
arXiv:2608. 04766v1 Announce Type: cross Abstract: A large number of infants with congenital anomalies are born each year globally, especially in areas with underdeveloped medical resources.
By Bin Pu, Jiewen Yang, Liwen Wang, Ying Tan, Guannan He, Xingbo Dong, Qika Lin, Jiarong Guo, Lixian Yang, Zuozhu Liu, Shengli Li, Kenli Li
arXiv:2606. 11106v1 Announce Type: cross Abstract: A global shortage of trained sonographers limits prenatal ultrasound screening in low- and middle-income countries, where over half of pregnant women receive no skilled sonography.
By Mahmood Alzubaidi, Uzair Shah, Raden Muaz, Ines Abbes, Nader Mohammed, Abdullatif Magram, Khalid Alyafei, Mowafa Househ, Marco Agus
arXiv:2608. 14763v1 Announce Type: cross Abstract: Assessment of ventriculomegaly (VM) on fetal brain ultrasound relies primarily on measuring lateral ventricular atrial width on standard planes, which is operator-dependent and may not fully reflect the overall ventricular enlargement.
By Yuhao Huang, Yuanji Zhang, Yuhuan Lu, Dong Ni, P. Ellen Grant, Davood Karimi
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
arXiv:2606. 29586v1 Announce Type: cross Abstract: Vision-language foundation models have shown strong potential in medical image analysis.
By Hang Su, Chao Sun, Zhaofan Li, Wei Hu, Juhua Liu, Bo Du
arXiv:2609.19230v1 Announce Type: new
Abstract: Ultrasound is the most widely deployed imaging modality worldwide, yet clinical AI remains fragmented into narrow single-task models that fail when dev...
By Chao Qin, Fahad Shahbaz Khan, Salman Khan, Sarim Ather, Siddiq Anwar, Rao Muhammad Anwer, Shadab Khan
arXiv:2606. 04705v1 Announce Type: cross Abstract: Semantic segmentation in medical imaging is a critical yet challenging task due to data scarcity and high variability across modalities.
By Amirhossein Movahedisefat, Amirreza Fateh, Mohammad Reza Mohammadi