arXiv:2606. 17437v1 Announce Type: cross Abstract: Automated classification of standard echocardiographic views is crucial for efficient clinical workflow but faces three main challenges.
By Bo Gou, Jicheng Zhang, Jianlong Xiong, Tao He, Bentian Liu, Hai Wu, Yijiao Wang, Yu Zhang, Yujia Yang, Yun Dai, Jian Liu, Jie Wang
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.
By Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
arXiv:2606. 31198v1 Announce Type: cross Abstract: Real-time video segmentation of the prostate in Transrectal Ultrasound (TRUS) is essential for image-guided interventions.
By Dong Yeong Kim, JunGyu Lee, Jaewon Choi, June Young Seo, Myeongseop Kim, Jinwook Choi, Taek Min Kim, Young-Gon Kim
The paper introduces ExiL, a mask‑conditioned progressive learning framework for bone ultrasound segmentation that models annotation as a structured refinement trajectory. ExiL uses a synthetic expert‑like brush simulator and a lightweight U‑Net to learn from imperfect masks, and it can be updated in real time from expert refinements. In experiments on UltraBones100k and a prospective volunteer dataset, ExiL cut average annotation time from 60 to 20 seconds per frame and improved mean Dice by about 0.045, achieving 0.87 Dice and 2.7 px boundary error with 10–50 ms inference.
By Arash Tavangar, Larissa K. Chiu, Hamidreza Khodashenas, Gregory K. Berry, Amir Hooshiar
X‑LMC is a spatiotemporal deep‑learning framework that automatically scores leptomeningeal collateral (LMC) status from time‑resolved biplane digital subtraction angiography (DSA). It uses a DINOv2 backbone to encode spatial frames, a token‑level cross‑view attention module to fuse orthogonal projections, and a recurrent network to model contrast bolus dynamics. On a multicenter dataset of 134 M1‑segment occlusion patients, X‑LMC achieved a Quadratic Weighted Kappa of 0.398 and a macro‑F1 of 0.711, outperforming static and other spatiotemporal baselines and matching clinical inter‑rater agreement.