arXiv:2605. 25402v2 Announce Type: replace-cross Abstract: Self-supervised pre-training paradigm has gained increasing prominence for learning transferable representations in medical imaging, yet existing methods for ultrasound (US) images operate at the image or frame level, overlooking the anatomical context for clinical-aligned representation learning.
By Chunzheng Zhu, Yijun Wang, Jianxin Lin, Feng Wang, Hongwei Wang, Lei Zhao, Shengli Li, Kenli Li
arXiv:2607.10851v2 Announce Type: replace
Abstract: Medical image classification models are ideally expected to identify diagnostically relevant regions while making predictions, yet standard classif...
By Tonmoy Hossain, Atiqur Rahman, Farhana Hossain Swarnali, Miaomiao Zhang
DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.
By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
Medical image classification models are ideally expected to identify diagnostically relevant regions while making predictions, yet standard classification losses rarely provide spatial supervision. Explicit supervision via anatomical shape information, such as segmentation masks of task-relevant anatomy, has been shown to guide the network toward regions relevant to the target prediction.
This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.
By Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang