arXiv Computer Vision

SCOPE-4D: Endoscopic 4D Geometry Foundation Models

SCOPE-4D is an endoscopic 4D geometry foundation model that predicts camera parameters, dense geometry, and 3D tissue trajectories from monocular RGB video in a single forward pass. The authors introduce SCOPE-5K, a curated dataset of about 5,000 real and synthetic gastrointestinal endoscopy and laparoscopy clips, and use it for geometric supervised fine‑tuning. Adding Common–Residual Motion (CRM) constraints and trajectory supervision further improves camera and depth estimation and enables dense 3D tissue tracking, as shown by evaluations on public and new benchmarks and a blinded user study.

arXiv AI
Jun 17

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.

By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath
arXiv Computer Vision
Aug 26

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.

By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
arXiv Computer Vision
Sep 11

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.

By Xinwei Ju, Rema Daher, Razvan Caramalau, Baoru Huang, Danail Stoyanov, Francisco Vasconcelos
arXiv Computer Vision
Aug 25

From Articulated Kinematics to Routed Visual Control for Action-Conditioned Surgical Video Generation

arXiv:2605.08712v2 Announce Type: replace Abstract: Action-conditioned surgical video generation is a critical yet highly challenging problem for robotic surgery. The core difficulty is that low-dime...

By Bohan Li, Shuojue Yang, Baorui Peng, Xianda Guo, Erli Zhang, Youqi Tao, Junfeng Duan, Daguang Xu, Qi Dou, Xin Jin, Wenjun Zeng, Hao Zhao, Yueming Jin
arXiv AI
Jul 10

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

arXiv:2607. 08408v1 Announce Type: cross Abstract: Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy.

By Tianyi Song, Sierra Bonilla, Xinwei Ju, Evangelos Mazomenos, Danail Stoyanov, Adam Schmidt, Omid Mohareri, Sophia Bano, Francisco Vasconcelos
Hugging Face Trending Papers
Jul 9

Track2Map: Online Deformable SLAM with Motion-Aware Pose Optimization in Robotic Surgery

Gaussian splatting is the current state-of-the-art for dense, deformable 3D anatomy reconstruction in robot-assisted minimally invasive surgery (RAMIS); however, most pipelines are offline and depend on accurate camera trajectory priors (often from robotic kinematics), limiting applicability when priors are missing or noisy. To address these limitations, we propose Track2Map, an online 3D Gaussian Splatting pipeline that jointly optimizes camera trajectory and 3D deformable scene representation directly from surgical video.