arXiv Machine Learning By Ashwath Radhachandran, Adam Tupper, Christian Gagn\'e, William Speier

UltraBench 2: Towards Robust Evaluation of Vision Foundation Models on Ultrasound

Read the original on arXiv Machine Learning →

UltraBench 2 is a new benchmark designed to evaluate vision foundation models on ultrasound images, addressing the lack of standardized tests in this area. It covers a wide range of anatomical structures and tasks, emphasizing reproducibility and ease of use. The authors compare existing models, finding that ultrasound-specific pretraining still outperforms on classification, while general-purpose models have matched performance on segmentation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
Aug 27

UltraPIPS: Improving model perception in B-mode ultrasound with foundation models

UltraPIPS introduces domain‑specific foundation models for measuring perceptual similarity in B‑mode ultrasound images. The study shows that ultrasound‑trained LPIPS backbones better correlate with downstream tasks such as classification, segmentation, and reconstruction than natural‑image or general medical models. Optimizing LPIPS loss with an ultrasound backbone yields a strong balance between reconstruction quality and realism, and the authors provide an open‑source library for these metrics.

By Tal Grutman, Tali Ilovitsh
arXiv AI
Aug 25

SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging

The paper introduces Segment Anything Small (SAS), a data‑augmentation method that improves deep‑learning segmentation of small anatomical structures in ultrasound images. SAS uses two transformations: resizing and embedding organ thumbnails into a black background to vary organ scale, and adding noise to regions of interest to mimic tissue texture variability. Experiments on one internal and five external datasets show Dice score gains up to 0.35, with an average improvement of 0.16, and demonstrate that SAS enhances model robustness and generalizability without adding hallucinations or artifacts.

By Danielle L. Ferreira, Ahana Gangopadhyay, Hsi-Ming Chang, Ravi Soni, Gopal Avinash
arXiv AI
Sep 10

US-JEPA: A Joint Embedding Predictive Architecture for Ultrasound

US-JEPA introduces a self‑supervised framework for ultrasound imaging that predicts masked latent representations instead of raw pixels, using a frozen, domain‑specific teacher to provide stable targets. This approach avoids the hyperparameter sensitivity and computational cost of traditional online teachers, enabling the student model to build upon the teacher’s semantic priors. The authors benchmark US‑JEPA against all publicly available ultrasound foundation models on UltraBench, showing competitive or superior performance across multiple organs and pathological conditions under linear probing.

By Ashwath Radhachandran, Vedrana Ivezi\'c, Shreeram Athreya, Corey W. Arnold, William Speier
arXiv AI
Jun 16

Enabling Real-Time Point-of-Care Ultrasound Segmentation: A GPU-Free Deployment in Resource-Limited Settings

arXiv:2606. 15176v1 Announce Type: cross Abstract: Ultrasound imaging is the most widely adopted medical modality globally due to its low cost and portability, yet artificial intelligence (AI) deployment remains constrained by reliance on GPU-accelerated models, creating a structural paradox where the cost of "intelligence" exceeds that of the imaging device itself.

By Weihao Gao
arXiv AI
6d ago

UltraG-Bench: A Multi-task Benchmark for assessing Large Vision-Language Models on Pixel-level Evidence Grounding in Ultrasound

UltraG-Bench is a large‑scale, multi‑task benchmark designed to evaluate pixel‑level evidence grounding in ultrasound images. It comprises 40 public segmentation datasets covering 13 anatomical categories and includes three progressive tasks—instruction‑guided segmentation, evidence‑grounded VQA, and evidence‑grounded report generation—with a total of 736,726 annotations. Evaluation of 14 state‑of‑the‑art models shows a significant gap between semantic understanding and fine‑grained pixel‑level localization, and the authors propose UltraG‑Agent, which combines a vision‑language model with the ultrasound‑specific segmentation model UltraSAM to improve both semantic prediction and visual grounding.

By Quanhao Zhu, Bo Xu, Rui Lin, Chenyuan Wang, Yu Shao, Boling Zhu, Jiuyan Sun, Liang Zhao, Hongfei Lin, Feng Xia
arXiv Computer Vision
Sep 2

Expert-like Bone Ultrasound Segmentation through Expert-in-the-loop Mask-conditioned Progressive Learning

The paper introduces ExiL, a mask‑conditioned progressive learning framework for bone ultrasound segmentation that models annotation as a structured refinement trajectory. ExiL uses a synthetic expert‑like brush simulator and a lightweight U‑Net to learn from imperfect masks, and it can be updated in real time from expert refinements. In experiments on UltraBones100k and a prospective volunteer dataset, ExiL cut average annotation time from 60 to 20 seconds per frame and improved mean Dice by about 0.045, achieving 0.87 Dice and 2.7 px boundary error with 10–50 ms inference.

By Arash Tavangar, Larissa K. Chiu, Hamidreza Khodashenas, Gregory K. Berry, Amir Hooshiar