arXiv Computer Vision

Lesion-centered 3D mapping of colonoscopy procedures: validation of a hierarchical ensemble pipeline on public benchmark videos

The paper presents a four‑layer hierarchical pipeline that constructs a lesion‑centered spatial record from colonoscopy videos without full‑colon 3D reconstruction. It combines a global topological map, lesion‑level spatio‑temporal tracks, on‑demand local 3D reconstruction, and persistent lesion identity across repeated observations, and evaluates the system on four public videos. The results show successful detection of revisit events, accurate lesion identity merging, and superior geometry accuracy compared to a general‑purpose foundation model.

arXiv Computer Vision
Aug 26

C3VDReg: A Benchmark for Local-to-Local Colonoscopic Registration toward Anatomical Localization

C3VDReg is a benchmark for local-to-local colonoscopic registration that uses the Colonoscopy 3D Video Dataset (C3VD) to generate 10,015 partial-to-partial point cloud pairs, with 2,088 held‑out test pairs. Each pair consists of a source point cloud from depth reprojection and a target point cloud from CT mesh raycasting, evaluated under a standardized protocol of 8,192 points per cloud and fixed pose conventions. Experiments show that high geometric overlap does not guarantee reliable pose recovery, revealing translation ambiguity along repetitive tubular anatomy as a key failure mode.

By Linzhe Jiang, Jiayuan Huang, Sophia Bano, Matthew J. Clarkson, Zhehua Mao, Mobarak I. Hoque
arXiv Computer Vision
Sep 11

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.

By Xinwei Ju, Rema Daher, Razvan Caramalau, Baoru Huang, Danail Stoyanov, Francisco Vasconcelos
Hugging Face Trending Papers
Jul 9

Metrics or Mirage? An Audit of Evaluation Inconsistencies in Colonoscopy Polyp Segmentation Benchmarks

Progress in colonoscopy polyp segmentation is routinely reported through leaderboard comparisons on a small set of public benchmarks. We argue that this apparent progress is difficult to verify: a systematic audit of \textbf{27 papers} published between 2015 and 2026 reveals three structural problems in how the community evaluates models.

arXiv Computer Vision
Sep 22

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.

By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv AI
Sep 10

WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos

WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.

By Giseong Hwang, Minjae Jo, Yeonghyeon Park, Kyeonghun Kim, Seoyeon Han, Donghoon Han, Haneul Kim, Yului Jeong, Insung Hwang, Pa Hong, Ken Ying-Kai Liao, Nam-Joon Kim
arXiv Computer Vision
Sep 18

RAUL: Reference-Assisted Ureteroscopy Localization for Skill Assessment

RAUL is a reference‑assisted reconstruction framework that recovers ureteroscope trajectories from endoscopic video alone, using a high‑quality reference exploration video for each phantom. It achieves a mean translation error of 0.5 mm and increases frame‑wise localization coverage from 50.5 % to 86.1 % compared to standard Structure‑from‑Motion. The reconstructed trajectories reveal significant differences in navigation metrics between high‑ and low‑experience trainees, enabling objective skill assessment without external tracking equipment.

By Fangjie Li, Mai Bui, Charan Mohan, Michael Miga, Matthieu Chabanas, Nicholas Kavoussi, Jie Ying Wu