arXiv Computer Vision

TTDF: A Two-Stage Framework for Reliable Surgical Phase Transition Detection

arXiv AI
Aug 19

SurgicalMamba: Dual-Path SSD with State Regramming for Online Surgical Phase Recognition

SurgicalMamba introduces a dual‑path Structured State‑Space Duality (SSD) model with two novel mechanisms—state regramming and intensity‑modulated stepping—to improve online surgical phase recognition. State regramming rotates the carried state at chunk boundaries based on content, separating repeated views that occur in different phases, while intensity‑modulated stepping adjusts decay rates at phase transitions to better capture phase length variability. The approach achieves state‑of‑the‑art accuracy and Jaccard scores on seven public benchmarks, running at 312.88 fps on a single GPU, and its rotation mechanism also boosts multi‑query associative recall in other domains.

By Sukju Oh, Sukkyu Sun
arXiv Computer Vision
Aug 25

Large-Small Model Collaboration for Zero-Shot Surgical Phase Recognition

The paper introduces LaST, a large-small collaborative framework for zero-shot surgical phase recognition. It combines a foundation model that generates frame-level phase priors with a lightweight model that refines predictions through iterative temporal refinement, dynamic quality control, and dual-model cross-learning. Experiments show LaST outperforms baseline and state-of-the-art methods, achieving significant accuracy gains on unseen clinical domains.

By Yiyi Zhang, Ying Zheng, Wenxin Fan, Yu Zhu, Yuchen Yuan, Litao Zhao, Zheng Li, Pheng-Ann Heng
arXiv Computer Vision
Sep 14

OphBiWSSD: Scaling Temporal Action Localization in Ophthalmic Surgeries with Bidirectional Weight-tied State Space Duality

OphBiWSSD is a new framework for temporal action localization in ophthalmic surgeries that uses Bidirectional State Space Duality to avoid the quadratic memory cost of Transformers. It employs a weight‑tied selective scan that incorporates both past and future surgical context, enabling linear‑time global synthesis of non‑causal temporal cues. On the OphNet benchmark, OphBiWSSD achieves state‑of‑the‑art mean Average Precisions of 44.42 % for phases and 43.08 % for operations, outperforming baselines by 6.80 % and 6.66 % respectively.

By Yang Liu, Qionghong Ma, Joongwon Chae, Lihui Luo, Yibing Shen, Yulin Zhuo, Yingting Zhu, Jiashu Chang, Xiaoyun Zhong, Dongmei Yu, Peter E. Lobie, Peiwu Qin, Chengming Yang
arXiv Computer Vision
Sep 18

State-Change Learning for Prediction of Future Events in Endoscopic Videos

The paper introduces SurgFUTR, a state‑change learning framework for predicting future events in endoscopic videos. Instead of forecasting raw observations, it classifies transitions between current and future states using a teacher‑student architecture and an Action Dynamics module. The authors also present SFPBench, a benchmark with five short‑ and long‑term prediction tasks, and demonstrate consistent improvements across multiple datasets and procedures, including cross‑procedure transfer.

By Saurav Sharma, Chinedu Innocent Nwoye, Didier Mutter, Nicolas Padoy
arXiv Computer Vision
Aug 21

Artificial Intelligence for Workflow Analysis in Colorectal Surgery: A Multicentric, Cross-Procedural Development and Generalization Study

arXiv:2608. 20154v1 Announce Type: new Abstract: Minimally invasive colorectal surgeries (MIS-CRS) are characterised by significant variability and inconsistent outcomes.

By Pietro Mascagni, Julia Alekseenko, Pooja P Jain, Marta Goglia, Andrea Balla, Ludovica Baldari, Gianfranco Silecchia, Claudio Fiorillo, Vincenzo Tondolo, Salvador Morales-Conde, Luigi Boni, Sergio Alfieri, Nicolas Padoy
Hugging Face Trending Papers
Jun 25

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.

arXiv Computer Vision
Sep 16

TEDi: Temporal Memory-Enhanced and Denoising Transformer for Surgical Instrument Segmentation

TEDi is a Temporal memory-Enhanced and Denoising Transformer designed for surgical instrument segmentation. It introduces a query-level memory bank with a memory search enhancement encoder to incorporate discriminative representations from past frames, and a temporal consistency denoising module that builds a cross‑frame semantic anchor to stabilize predictions. Experiments on EndoVis 2017 and EndoVis 2018 show that TEDi outperforms existing state‑of‑the‑art methods, indicating its effectiveness for computer‑assisted surgery.

By Jiahong Yuan, Weiming Mi, Tao Zhang, Haoyin Zhou