arXiv AI

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.

arXiv Computer Vision
Aug 28

DALE-CT: Depth-Aware 2D Slice Encoders Learn an Anatomical World Model of Chest CT

DALE-CT introduces depth‑aware 2D slice encoders that learn an anatomical world model of chest CT scans without 3D or positional supervision. By sampling self‑supervised views across a physical $z$‑axis slab, the encoder captures how anatomy changes between neighboring slices, enabling it to recover slice ordering and distinguish slices by anatomy alone. The model, trained on a large 287k‑scan corpus, achieves state‑of‑the‑art performance on CT‑RATE and is released with full code and evaluation tools.

By Evan W. Damron, Mahmut S. Gokmen, Mitchell A. Klusty, Caroline N. Leach, Emily B. Collier, V. K. Cody Bumgardner
arXiv AI
Sep 10

US-JEPA: A Joint Embedding Predictive Architecture for Ultrasound

US-JEPA introduces a self‑supervised framework for ultrasound imaging that predicts masked latent representations instead of raw pixels, using a frozen, domain‑specific teacher to provide stable targets. This approach avoids the hyperparameter sensitivity and computational cost of traditional online teachers, enabling the student model to build upon the teacher’s semantic priors. The authors benchmark US‑JEPA against all publicly available ultrasound foundation models on UltraBench, showing competitive or superior performance across multiple organs and pathological conditions under linear probing.

By Ashwath Radhachandran, Vedrana Ivezi\'c, Shreeram Athreya, Corey W. Arnold, William Speier
arXiv AI
Sep 18

Optimal Transport Metric Learning for Feature Alignment in Partially Supervised Segmentation

The paper proposes a two‑stage learning framework for multi‑organ segmentation that handles partially annotated datasets and domain shifts. First, the model learns accurate segmentations from available annotations to build robust feature representations. Second, it introduces learnable organ prototypes and a Sinkhorn‑triplet loss to enforce organ‑wise feature consistency across datasets, keeping embeddings of the same organ close while separating different organs, even when annotations are missing.

By Dakini Mallam Garba, Salim Abdou Daoura
Hugging Face Trending Papers
Jul 12

Learning To Focus: Anatomy-Guided Attention Regularization for Medical Image Classification

Medical image classification models are ideally expected to identify diagnostically relevant regions while making predictions, yet standard classification losses rarely provide spatial supervision. Explicit supervision via anatomical shape information, such as segmentation masks of task-relevant anatomy, has been shown to guide the network toward regions relevant to the target prediction.

arXiv AI
Aug 25

SAS: Segment Anything Small for Ultrasound -- A Non-Generative Data Augmentation Technique for Robust Deep Learning in Ultrasound Imaging

The paper introduces Segment Anything Small (SAS), a data‑augmentation method that improves deep‑learning segmentation of small anatomical structures in ultrasound images. SAS uses two transformations: resizing and embedding organ thumbnails into a black background to vary organ scale, and adding noise to regions of interest to mimic tissue texture variability. Experiments on one internal and five external datasets show Dice score gains up to 0.35, with an average improvement of 0.16, and demonstrate that SAS enhances model robustness and generalizability without adding hallucinations or artifacts.

By Danielle L. Ferreira, Ahana Gangopadhyay, Hsi-Ming Chang, Ravi Soni, Gopal Avinash
arXiv Machine Learning
Sep 14

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

DenseTRF is a self‑supervised framework that adapts texture‑aware representations for dense prediction in surgical computer vision. It uses slot attention to learn invariant visual structures and then conditions dense prediction on these representations, merging models to adapt to target distributions without supervision. Experiments on multiple surgical procedures show that DenseTRF improves cross‑distribution generalization compared to state‑of‑the‑art segmentation models and test‑distribution adaptation methods.

By Guiqiu Liao, Matja\v{z} Jogan, Daniel A. Hashimoto
arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri