arXiv AI

Vis2Reg: Visibility-Aware Landmark-Free Geometric 3D--2D Registration for Liver Laparoscopy

arXiv:2607. 17810v1 Announce Type: cross Abstract: Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision.

arXiv Computer Vision
Sep 18

LapaTrack-3D: 6 DoF pre-operative shape tracking for laparoscopic surgery

LapaTrack-3D is a real‑time 6 DoF tracking algorithm designed for monocular laparoscopic surgery, aligning intra‑operative video with pre‑operative CT data. It modifies the ORB‑SLAM2 framework with four key changes: fast initialization using a primitive 3D shape, a pseudo‑segmentation strategy to isolate the target organ, incorporation of the 3D shape as a geometric prior in pose graph optimization, and enhanced imaging via a modified Multi‑Scale Retinex with Chromaticity Preservation algorithm. Experiments on in‑vivo and ex‑vivo data show robust tracking under challenging conditions such as poor illumination, fast motion, and partial visibility, achieving 13 Hz on 1280×720 video.

By Jingwei Song, Javid Hussain Jakir, Ray Zhang, Wenwei Zhang, Hao Zhou, Xiaomeng Xian, Maani Ghaffari
arXiv Computer Vision
Aug 25

Dense Structural Priors for Sparse Functional Landmark Localization in Surgical Videos

The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.

By Chenyan Jing, Hao Ding, Lalithkumar Seenivasan, Jacob M. Delgado L\'opez, Mathias Unberath
arXiv Computer Vision
6d ago

PICO: Projection-Informed Consistency Optimisation for 6DoF Surgical Tool Pose Estimation

The paper introduces PICO, an end-to-end trainable model for 6DoF surgical tool pose estimation that uses multi-task learning to predict segmentation, depth, and pose parameters. It incorporates two geometry-aware proxy tasks—a projection loss and a point-to-point loss—to enforce consistency in 2D and 3D spaces, improving accuracy and robustness. Evaluated on the SurgRIPE dataset, PICO achieves strong performance, ranking second in rotation accuracy and maintaining competitive translation results, especially under occlusion.

By Lucy Fothergill, Pietro Valdastri, Dominic Jones, Duygu Sarikaya
arXiv Computer Vision
Aug 26

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.

By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
arXiv Computer Vision
Aug 28

Test Time Adaptation Methods for Point Cloud Registration in Laparoscopic Surgery

The paper investigates test‑time adaptation (TTA) techniques for 3D point‑cloud registration in laparoscopic surgery, where synthetic training data must be adapted to noisy, sparse, and occluded real intraoperative reconstructions. It adapts three families of TTA methods—model, normalization, and input adaptation—to handle asymmetric shifts between preoperative meshes and intraoperative clouds, replacing classification‑based entropy objectives with correspondence‑based ones. Experiments on synthetic and real targets show that input adaptation consistently reduces registration error with low inference latency, making it the most promising approach for surgical applications.

By Nina Bodelot, Soufiane Belharbi, Eric Granger
arXiv AI
Jun 17

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.

By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath