arXiv Computer Vision

Match4Annotate: Cross-Video Annotation Transfer in Ultrasound via Implicit Feature Flow-Guided Matching

Match4Annotate is a test‑time framework that transfers user‑specified annotations from a labeled ultrasound video to an unlabeled target video without requiring manual initialization. It uses a spatiotemporal implicit feature representation, a continuous implicit feature flow for alignment, and flow‑guided annotation transfer to unify sparse point and dense mask transfer. The method achieves state‑of‑the‑art performance on four clinical ultrasound datasets, outperforming dense feature‑matching baselines and one‑shot segmentation methods, and works without task‑specific training on a single consumer GPU.

arXiv Computer Vision
Sep 22

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.

By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv AI
Jun 3

Anatomy-Anchored Self-Supervision: Distilling Vision Foundation Models for Invariant Ultrasound Representation

arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.

By Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
arXiv Computer Vision
Sep 2

Expert-like Bone Ultrasound Segmentation through Expert-in-the-loop Mask-conditioned Progressive Learning

The paper introduces ExiL, a mask‑conditioned progressive learning framework for bone ultrasound segmentation that models annotation as a structured refinement trajectory. ExiL uses a synthetic expert‑like brush simulator and a lightweight U‑Net to learn from imperfect masks, and it can be updated in real time from expert refinements. In experiments on UltraBones100k and a prospective volunteer dataset, ExiL cut average annotation time from 60 to 20 seconds per frame and improved mean Dice by about 0.045, achieving 0.87 Dice and 2.7 px boundary error with 10–50 ms inference.

By Arash Tavangar, Larissa K. Chiu, Hamidreza Khodashenas, Gregory K. Berry, Amir Hooshiar
Hugging Face Trending Papers
Aug 19

X-LMC: Cross-View Spatiotemporal Collateral Circulation Scoring from DSA

X‑LMC is a spatiotemporal deep‑learning framework that automatically scores leptomeningeal collateral (LMC) status from time‑resolved biplane digital subtraction angiography (DSA). It uses a DINOv2 backbone to encode spatial frames, a token‑level cross‑view attention module to fuse orthogonal projections, and a recurrent network to model contrast bolus dynamics. On a multicenter dataset of 134 M1‑segment occlusion patients, X‑LMC achieved a Quadratic Weighted Kappa of 0.398 and a macro‑F1 of 0.711, outperforming static and other spatiotemporal baselines and matching clinical inter‑rater agreement.

arXiv Computer Vision
Sep 7

DART: Depth-as-Target Pretraining for Surgical Vision Foundation Models

DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.

By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv Computer Vision
Sep 15

Echo-E$^3$Net: Efficient Endocardial Spatio-Temporal Network for Ejection Fraction Estimation

Echo-E$^3$Net is an anatomy‑guided spatio‑temporal neural network designed to estimate left ventricular ejection fraction (LVEF) from ultrasound images. It uses a dual‑phase Endocardial Border Detector to locate end‑diastole and end‑systole landmarks and an Endocardial Feature Aggregator to fuse these landmarks with global deep‑feature descriptors for EF regression. The model achieves competitive accuracy on EchoNet‑Dynamic and EchoNet‑Pediatric datasets while using only 1.55 M parameters and 8.05 GFLOPs, enabling real‑time deployment on limited‑resource devices.

By Moein Heidari, Afshin Bozorgpour, AmirHossein Zarif-Fakharnia, Wenjin Chen, Dorit Merhof, David J. Foran, Jasmine Grewal, Ilker Hacihaliloglu
Hugging Face Trending Papers
Jun 25

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.

arXiv Computer Vision
Aug 28

Unsupervised Adaptation of 3D CT Foundation Models for 3D CBCT Segmentation

The paper introduces an unsupervised domain adaptation framework that aligns redundancy-reducing features to enable accurate 3D segmentation of cone-beam CT (CBCT) without target-domain annotations or inference-time adaptation. The method is architecture-agnostic, working with both CNN-based and ViT-based foundation models, and is evaluated on two liver segmentation benchmarks for interventional vascular procedures and radiation therapy. Results show that even large pretrained segmentation networks need explicit feature-space bridging to generalize across diagnostic CT and CBCT, and the proposed approach consistently outperforms existing pretrained foundation models and UDA strategies.

By Gauthier Miralles, Loic Le Folgoc, Vincent Jugnon, Pietro Gori