arXiv Machine Learning

S3-Tracker: Self-Supervised Surgical Tissue Tracking With Contrastive Random Walks

Hugging Face Trending Papers
Jul 9

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy. To address these limitations, we propose Track2Map, an online 3D Gaussian Splatting pipeline that jointly optimizes camera trajectory and 3D deformable scene representation directly from surgical video.

arXiv AI
Jul 10

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

arXiv:2607. 08408v1 Announce Type: cross Abstract: Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy.

By Tianyi Song, Sierra Bonilla, Xinwei Ju, Evangelos Mazomenos, Danail Stoyanov, Adam Schmidt, Omid Mohareri, Sophia Bano, Francisco Vasconcelos
arXiv AI
Jul 21

Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy

arXiv:2607. 17810v1 Announce Type: cross Abstract: Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision.

By Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali
arXiv Computer Vision
Aug 25

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

arXiv:2605.08712v2 Announce Type: replace Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dime...

By Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin
Hugging Face Trending Papers
Jun 25

Temporally Consistent Label Interpolation for Robust Surgical Multi-Task Learning under Challenging Conditions

Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.

arXiv Computer Vision
Sep 18

LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

LapaTrack-3D is a real‑time 6 DoF tracking algorithm designed for monocular laparoscopic surgery, aligning intra‑operative video with pre‑operative CT data. It modifies the ORB‑SLAM2 framework with four key changes: fast initialization using a primitive 3D shape, a pseudo‑segmentation strategy to isolate the target organ, incorporation of the 3D shape as a geometric prior in pose graph optimization, and enhanced imaging via a modified Multi‑Scale Retinex with Chromaticity Preservation algorithm. Experiments on in‑vivo and ex‑vivo data show robust tracking under challenging conditions such as poor illumination, fast motion, and partial visibility, achieving 13 Hz on 1280×720 video.

By Jingwei Song, Javid Hussain Jakir, Ray Zhang, Wenwei Zhang, Hao Zhou, Xiaomeng Xian, Maani Ghaffari
arXiv Computer Vision
Aug 28

Surgical Video Generation From Diffusion to World Models: A Survey

This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.

By Fuxiang Huang, Chenxu Zhang, Liang Han, Lei Zhang
arXiv Computer Vision
Sep 24

Track2Art: Motion-Centric Articulated Object Model Recovery from 2D Point Trackers

Track2Art is a motion‑centric framework that recovers articulated object models from RGB‑D interaction videos by lifting 2D point tracks into 3D trajectories. It groups these trajectories into rigid‑part hypotheses and uses learned‑analytic reasoning to infer directed kinematic relations, joint types, and joint geometry. On the PartNet‑Mobility benchmark, it achieves 0.695 Point IoU and 0.410 end‑to‑end J@20 without requiring ground‑truth part counts or test‑time optimization.

By Xiaotong Li, Yixiong Jing, Junsheng Ding, Weihang Li, Benjamin Busam, Guangming Wang, Brian Sheil
arXiv Machine Learning
Sep 14

DenseTRF: Texture-Aware Unsupervised Representation Adaptation for Surgical Scene Dense Prediction

DenseTRF is a self‑supervised framework that adapts texture‑aware representations for dense prediction in surgical computer vision. It uses slot attention to learn invariant visual structures and then conditions dense prediction on these representations, merging models to adapt to target distributions without supervision. Experiments on multiple surgical procedures show that DenseTRF improves cross‑distribution generalization compared to state‑of‑the‑art segmentation models and test‑distribution adaptation methods.

By Guiqiu Liao, Matja\v{z} Jogan, Daniel A. Hashimoto