arXiv Computer Vision

LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

LapaTrack-3D is a real‑time 6 DoF tracking algorithm designed for monocular laparoscopic surgery, aligning intra‑operative video with pre‑operative CT data. It modifies the ORB‑SLAM2 framework with four key changes: fast initialization using a primitive 3D shape, a pseudo‑segmentation strategy to isolate the target organ, incorporation of the 3D shape as a geometric prior in pose graph optimization, and enhanced imaging via a modified Multi‑Scale Retinex with Chromaticity Preservation algorithm. Experiments on in‑vivo and ex‑vivo data show robust tracking under challenging conditions such as poor illumination, fast motion, and partial visibility, achieving 13 Hz on 1280×720 video.

arXiv AI
Jul 21

Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy

arXiv:2607. 17810v1 Announce Type: cross Abstract: Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision.

By Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali
Hugging Face Trending Papers
Jul 9

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy. To address these limitations, we propose Track2Map, an online 3D Gaussian Splatting pipeline that jointly optimizes camera trajectory and 3D deformable scene representation directly from surgical video.

arXiv AI
Jul 10

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

arXiv:2607. 08408v1 Announce Type: cross Abstract: Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy.

By Tianyi Song, Sierra Bonilla, Xinwei Ju, Evangelos Mazomenos, Danail Stoyanov, Adam Schmidt, Omid Mohareri, Sophia Bano, Francisco Vasconcelos
arXiv Computer Vision
Sep 2

Online camera-pose-free stereo endoscopic tissue deformation recovery with tissue-invariant vision-biomechanics consistency

The paper presents a camera‑pose‑free stereo endoscopic method for recovering tissue deformation by modeling geometry as a 3D point‑derivative map and deformation as a 3D displacement‑local deformation map. It optimizes inter‑frame deformation in a camera‑centric setting, eliminating the need for camera pose estimation, and introduces a canonical map for online geometry and deformation optimization. Experiments on in‑vivo and ex‑vivo laparoscopic data show accurate 3D reconstruction (≈0.37–0.39 mm surface distance) even under occlusion, and the method can estimate surface strain distributions during manipulation.

By Jiahe Chen, Naoki Tomii, Ichiro Sakuma, Etsuko Kobayashi
arXiv Computer Vision
Sep 3

MV-dVRK: A Multi-Viewpoint Benchmark for Spatial Surgical Perception

MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.

By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
arXiv AI
Aug 10

Surg-UniWorld: A Unified Surgical World Model with Multimodal Control Experts

arXiv:2608. 06770v1 Announce Type: new Abstract: Controllable surgical world models can provide a generative foundation for surgical artificial intelligence and simulation by synthesizing realistic instrument--tissue interactions.

By Rulin Zhou, Wanhao Liu, Guoheng Ma, Liangjin Shao, Qiujie Song, Yidu Wang, Guankun Wang, Tong Chen, Long Bai, Luping Zhou, Hongliang Ren
arXiv Computer Vision
4d ago

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.

By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
arXiv Computer Vision
Aug 25

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.

By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath