Surgical procedures follow a phase-to-step hierarchy, yet the video-language models used to recognize them are evaluated with flat per-level metrics that ignore cross-level coherence and error structu...
arXiv:2606. 29247v1 Announce Type: new Abstract: Vision-Language-Action (VLA) models represent a promising direction for embodied intelligence in surgical robotics.
By Jiashuo Sun, Yue He, Wenxuan Liu, Tao Mao, Jiazheng Wang, Xiang Chen, Min Liu
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
SurgAtlas is the largest surgical video‑language dataset, containing 15,291 videos (2,391 hours) across 18 specialties and over 5,000 procedure types, all sourced from public YouTube. It uniquely includes open‑surgery videos at scale (6,182) alongside more than 9,000 minimally invasive recordings, and introduces standardized benchmarks for open‑surgery video understanding. The dataset offers a rich, multi‑tier annotation schema—segment‑level captions, step/phase descriptions, video‑level surgical narratives, and reasoning‑oriented VQA pairs—validated by experts and built through an automated LLM‑enriched pipeline.
"whyItMatters":"SurgAtlas provides an unprecedentedly large, diverse, and clinically validated resource that can train and benchmark multimodal surgical AI models, advancing the development of next‑generation foundation models for surgery."
By Filippos Bellos, Andre S. Gala-Garza, Miaowei Wang, Alyssa M. Hardin, Ahmad M. Hider, Li Yayuan, Jing Bi, Susan Liang, Chenliang Xu, Donald S. Likosky, Jason J. Corso
We introduce SurgAtlas, the largest surgical video-language dataset to date, comprising 15,291 videos (2,391 hours) spanning 18 surgical specialties and over 5,000 procedure types, sourced entirely from publicly available YouTube content. SurgAtlas is also the first surgical video-language dataset to include open surgery at scale, with 6,182 open procedure videos alongside over 9,000 minimally invasive recordings, and the first to establish standardized benchmarks for open-surgery video understanding.
arXiv:2603.29962v4 Announce Type: replace
Abstract: Surgical procedures are inherently complex and risky, requiring extensive expertise and constant focus to navigate evolving intraoperative scenes....
By Shi Li, Vinkle Srivastav, Nicolas Chanel, Saurav Sharma, Nabani Banik, Lorenzo Arboit, Kun Yuan, Pietro Mascagni, Nicolas Padoy
arXiv:2608.24671v1 Announce Type: new
Abstract: Referring surgical video segmentation requires segmenting a target instrument or tissue region across video frames according to a natural language expr...
By Jiaxin Wen, Ming Yin, Lu Liu, Zeyu Fu
arXiv:2608.30872v1 Announce Type: new
Abstract: Objective assessment of surgical technical skill is important for surgical training and structured feedback, but current workflows remain dependent on...
By Chaohui Dang, Zheheng Jiang, James Glasbey, David Luke, Theodoros Arvanitis, Le Zhang
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
arXiv:2608. 06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions.
By Rulin Zhou, Wanhao Liu, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren
OphBiWSSD is a new framework for temporal action localization in ophthalmic surgeries that uses Bidirectional State Space Duality to avoid the quadratic memory cost of Transformers. It employs a weight‑tied selective scan that incorporates both past and future surgical context, enabling linear‑time global synthesis of non‑causal temporal cues. On the OphNet benchmark, OphBiWSSD achieves state‑of‑the‑art mean Average Precisions of 44.42 % for phases and 43.08 % for operations, outperforming baselines by 6.80 % and 6.66 % respectively.
By Yang Liu, Qionghong Ma, Joongwon Chae, Lihui Luo, Yibing Shen, Yulin Zhuo, Yingting Zhu, Jiashu Chang, Xiaoyun Zhong, Dongmei Yu, Peter E. Lobie, Peiwu Qin, Chengming Yang
The paper introduces SurgFUTR, a state‑change learning framework for predicting future events in endoscopic videos. Instead of forecasting raw observations, it classifies transitions between current and future states using a teacher‑student architecture and an Action Dynamics module. The authors also present SFPBench, a benchmark with five short‑ and long‑term prediction tasks, and demonstrate consistent improvements across multiple datasets and procedures, including cross‑procedure transfer.
By Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy