arXiv:2609.14624v1 Announce Type: new
Abstract: Reliable workflow transition detection is important for context-aware surgical assistance and downstream decision support. However, online surgical pha...
By Yushi Guo, Pietro Valdastri, Duygu Sarikaya
arXiv:2609.18971v1 Announce Type: new
Abstract: Surgical phase recognition maps each video frame to a clinically meaningful workflow phase, supporting context-aware assistance, documentation, and pos...
By Ye Tao, Claudia Scherl, Sara Monji-Azad
OphBiWSSD is a new framework for temporal action localization in ophthalmic surgeries that uses Bidirectional State Space Duality to avoid the quadratic memory cost of Transformers. It employs a weight‑tied selective scan that incorporates both past and future surgical context, enabling linear‑time global synthesis of non‑causal temporal cues. On the OphNet benchmark, OphBiWSSD achieves state‑of‑the‑art mean Average Precisions of 44.42 % for phases and 43.08 % for operations, outperforming baselines by 6.80 % and 6.66 % respectively.
By Yang Liu, Qionghong Ma, Joongwon Chae, Lihui Luo, Yibing Shen, Yulin Zhuo, Yingting Zhu, Jiashu Chang, Xiaoyun Zhong, Dongmei Yu, Peter E. Lobie, Peiwu Qin, Chengming Yang
The paper introduces SurgFUTR, a state‑change learning framework for predicting future events in endoscopic videos. Instead of forecasting raw observations, it classifies transitions between current and future states using a teacher‑student architecture and an Action Dynamics module. The authors also present SFPBench, a benchmark with five short‑ and long‑term prediction tasks, and demonstrate consistent improvements across multiple datasets and procedures, including cross‑procedure transfer.
By Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
The paper introduces LaST, a large-small collaborative framework for zero-shot surgical phase recognition. It combines a foundation model that generates frame-level phase priors with a lightweight model that refines predictions through iterative temporal refinement, dynamic quality control, and dual-model cross-learning. Experiments show LaST outperforms baseline and state-of-the-art methods, achieving significant accuracy gains on unseen clinical domains.
By Yiyi Zhang, Ying Zheng, Wenxin Fan, Yu Zhu, Yuchen Yuan, Litao Zhao, Zheng Li, Pheng-Ann Heng