Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue mo...
C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.
By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.
By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath
We present FoundationGeo, a two-stage framework that explicitly bridges relative and metric prediction via spatial calibration and principled data design. Stage 1 learns a high-fidelity, affine-invariant geometry model by initializing with DINOv3 and training on a curated 10.
The paper presents a four‑layer hierarchical pipeline that constructs a lesion‑centered spatial record from colonoscopy videos without full‑colon 3D reconstruction. It combines a global topological map, lesion‑level spatio‑temporal tracks, on‑demand local 3D reconstruction, and persistent lesion identity across repeated observations, and evaluates the system on four public videos. The results show successful detection of revisit events, accurate lesion identity merging, and superior geometry accuracy compared to a general‑purpose foundation model.
By Hyunjun Kim, Hyeonwoo Na, Jaewoo Lee
MV-dVRK is the first ex‑vivo surgical dataset that provides multiple exposure‑synchronized stereo viewpoints, accurate surface geometry, and ground‑truth camera poses for endoscopic images. The benchmark’s static subset offers dense SfM reference geometry validated against an industrial 3D scanner, while the dynamic sequences cover ten surgical tasks with increasing kinematic complexity and tissue deformation. Using MV‑dVRK, the authors systematically compare zero‑shot monocular, stereo, multi‑stereo, and multi‑view 3D reconstruction methods, finding that multi‑stereo reconstruction with two endoscopes yields the highest coverage, and that optimization‑based multi‑view methods outperform feed‑forward foundation models when a third viewpoint is added.
By Guido Caccianiga, Sergey Prokudin, Yutong Chen, Bernard Javot, Rachael L'Orsa, Omer Burak Alada\u{g}, Yarden Sharon, Jens Rolinger, Ivan Capobianco, Anton Deguet, Siyu Tang, Katherine J. Kuchenbecker
SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.
By Xinwei Ju, Rema Daher, Razvan Caramalau, Baoru Huang, Danail Stoyanov, Francisco Vasconcelos
arXiv:2607. 17810v1 Announce Type: cross Abstract: Accurate 3D--2D liver registration, which aligns preoperative 3D models to partial, view-dependent intraoperative surface observations, is critical for AR-guided laparoscopic surgery but remains challenging due to severe occlusion, limited visibility, and the lack of 3D ground-truth supervision.
By Jiaming Feng, Xukun Zhang, Shahid Farid, Sharib Ali
arXiv:2609.14313v1 Announce Type: cross
Abstract: Robust point tracking in endoscopic videos is essential for computer-assisted intervention and autonomous robotic surgery, enabling continuous regist...
By Jiaming Zhang, Zijian Wu, Mehran Armand, Septimiu Salcudean
arXiv:2607. 23343v1 Announce Type: cross Abstract: Intraoperative 2D/3D registration aligns preoperative CT volumes with intraoperative X-ray or fluoroscopic images and is essential for image-guided interventions.
By Minheng Chen, Youyong Kong
DART is a new RGB‑D pretraining method for surgical vision foundation models that incorporates pseudo‑labeled depth maps as a pixel‑space reconstruction target during training. By adding a depth reconstruction head to DINOv2’s masked iBOT framework, DART improves representation quality without affecting downstream RGB‑only fine‑tuning or inference. Across eight surgical benchmarks—including segmentation, depth estimation, and image‑level recognition—DART outperforms both natural‑image and in‑domain baselines, demonstrating that geometric pseudo‑labels can strengthen foundation model pretraining without extra labels or inference cost.
By John J. Han, Adam Schmidt, Muhammad Abdullah Jamal, Jie Ying Wu, Omid Mohareri
arXiv:2411. 17790v3 Announce Type: replace-cross Abstract: Accurate 3D mapping in endoscopy enables quantitative, holistic lesion characterization within the gastrointestinal (GI) tract, requiring reliable depth and pose estimation.
By Ziang Xu, Bin Li, Yang Hu, Chenyu Zhang, James East, Sharib Ali, Jens Rittscher