arXiv Computer Vision
4d ago

EndoPrior-GS: Dynamic Endoscopic Reconstruction with a Joint Texture Prior

EndoPrior-GS is a new pipeline for dynamic endoscopic reconstruction that combines frame-extracted vision heuristics with depth maps. It creates a joint texture prior using a tool-filtered tissue mask, a non-specular photometric filter, and anatomical salience, which guides primitive initialization and density control. Experiments on EndoNeRF and SCARED datasets show that EndoPrior-GS reduces Flow Error by 27.7% and 25.8% compared to representative methods while maintaining real-time rendering speed and competitive quality.

By Jiaqi Huang, Shidong Wang, Tong Xin, Kabita Adhikari
arXiv AI
Aug 10

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

arXiv:2608. 06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions.

By Rulin Zhou, Wanhao Liu, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren
arXiv Computer Vision
Sep 22

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.

By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv Computer Vision
Aug 28

Surgical Video Generation From Diffusion to World Models: A Survey

This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.

By Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang
arXiv Computer Vision
Sep 2

SurgiATM: A Physics-Guided Plug-and-Play Model for Deep Learning-Based Smoke Removal in Laparoscopic Surgery

The paper introduces SurgiATM, a lightweight physics-guided module for removing surgical smoke from laparoscopic endoscopic frames. It integrates a physics-based atmospheric model with a data-driven deep learning approach via a Mixture-of-Experts output stage, using a Laplacian-like error distribution to model smoke. SurgiATM adds only two hyperparameters and no extra trainable weights, enabling easy integration into existing desmoking architectures and improving accuracy and stability across multiple datasets and procedures.

By Mingyu Sheng, Jianan Fan, Dongnan Liu, Guoyan Zheng, Ron Kikinis, Weidong Cai
arXiv AI
Sep 10

Novel Methods for Catheter and Guidewire Segmentation in X-ray Fluoroscopy under a Federated Learning Setting

The paper introduces a structure‑aware federated learning framework for segmenting catheters and guidewires in X‑ray fluoroscopy. It presents a new benchmark dataset, CathAction, and a shape‑sensitive loss that improves segmentation accuracy. The approach extends to federated learning, adding projected gradient descent for adversarial optimization, and includes a diffusion‑based synthetic data generator that boosts performance under data scarcity.

By Chayun Kongtongvattana