SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline.
"whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."
By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso
arXiv:2608.24671v1 Announce Type: new
Abstract: Referring surgical video segmentation requires segmenting a target instrument or tissue region across video frames according to a natural language expr...
By Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.
arXiv:2609.21402v1 Announce Type: new
Abstract: Surgical instrument segmentation (SIS) plays a critical role in robotic assistance and surgical workflow analysis. However, most existing SIS methods f...
By Zhibo Zhang, Qijie Wang, Zengqiang Yan
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv:2608.24541v1 Announce Type: new
Abstract: Surgical instrument segmentation (SIS) is fundamental for computer-assisted surgery, where reliable instrument masks enable precise scene understanding...
By Xinning Yao, Jingjing Wang, Jinghua Yue, Xiaoyan Luo, Fugen Zhou, Bo Liu
arXiv:2607. 13237v1 Announce Type: cross Abstract: Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge.
By Manasa Dendukuri, Matjaz Jogan, Daniel A. Hashimoto, Guiqiu Liao
arXiv:2608.30872v1 Announce Type: new
Abstract: Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on...
By Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang
arXiv:2607.12896v3 Announce Type: replace
Abstract: Medical image segmentation foundation models are expected to generalize across diverse clinical scenarios, yet existing universal methods remain fr...
By Yunzhou Li, Jiesi Hu, Yanwu Yang, Hanyang Peng, Chenfei Ye, Jianfeng Cao, Yixuan Yuan, Ting Ma
Surgical instrument segmentation (SIS) is fundamental for computer-assisted surgery, where reliable instrument masks enable precise scene understanding and clinical assistance. Recently, adapting foun...
Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.