LapaTrack-3D is a real‑time 6 DoF tracking algorithm designed for monocular laparoscopic surgery, aligning intra‑operative video with pre‑operative CT data. It modifies the ORB‑SLAM2 framework with four key changes: fast initialization using a primitive 3D shape, a pseudo‑segmentation strategy to isolate the target organ, incorporation of the 3D shape as a geometric prior in pose graph optimization, and enhanced imaging via a modified Multi‑Scale Retinex with Chromaticity Preservation algorithm. Experiments on in‑vivo and ex‑vivo data show robust tracking under challenging conditions such as poor illumination, fast motion, and partial visibility, achieving 13 Hz on 1280×720 video.
By Jingwei Song, Javid Hussain Jakir, Ray Zhang, Wenwei Zhang, Hao Zhou, Xiaomeng Xian, Maani Ghaffari
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue mo...
arXiv:2609.23961v1 Announce Type: new
Abstract: Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflectio...
By Zhihao Xing, Yingyu Wang, Liang Zhao, Shoudong Huang
The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.
By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
arXiv:2609.14313v1 Announce Type: cross
Abstract: Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous regist...
By Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean
3D point cloud registration in laparoscopic surgery estimates the transformation between an intraoperative organ reconstructed from video and its preoperative mesh. Because ground-truth transformations are unavailable for real data, supervised networks are trained on synthetic organ pairs.
C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.
By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
The paper investigates test‑time adaptation (TTA) techniques for 3D point‑cloud registration in laparoscopic surgery, where synthetic training data must be adapted to noisy, sparse, and occluded real intraoperative reconstructions. It adapts three families of TTA methods—model, normalization, and input adaptation—to handle asymmetric shifts between preoperative meshes and intraoperative clouds, replacing classification‑based entropy objectives with correspondence‑based ones. Experiments on synthetic and real targets show that input adaptation consistently reduces registration error with low inference latency, making it the most promising approach for surgical applications.
By Nina Bodelot, Soufiane Belharbi, Eric Granger
arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.
By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath
arXiv:2606. 17379v1 Announce Type: cross Abstract: Accurate intraoperative liver registration is challenging due to substantial soft-tissue deformation yet sparse intraoperative measurements.
By Casey Meisenzahl, Jon Heiselman, Michael Holtz, Yubo Ye, Michael Miga, Linwei Wang
arXiv:2609.27227v1 Announce Type: new
Abstract: Objective assessment of robotic surgery uses instrument kinematics, which must be reconstructed when only video is available. We introduce a kinematic...
By Mehmet Kerem Turkcan, Soham Samal, Zoran Kostic