arXiv Computer Vision

SegCol Challenge: Semantic Segmentation for Tools and Fold Edges in Colonoscopy data

SegCol is a new dataset and benchmark for semantic segmentation of colon fold edges and surgical instruments in colonoscopy images, derived from the EndoMapper dataset. It offers manually annotated pixel‑level masks for three instrument classes and thin fold‑edge structures across temporally consistent image sequences, and serves as the basis for the SegCol Challenge within the EndoVis Challenge at MICCAI 2024. The study evaluates supervised segmentation and annotation‑efficient active learning, analyzes various segmentation metrics under structural perturbations, and highlights how metric behavior depends on target structure, underscoring the need for carefully selected evaluation protocols in endoscopic segmentation.

arXiv AI
1d ago

Colorectal Cancer Segmentation with Adaptive Augmentation and Multi-Resolution Ensemble Models

The paper presents an automated segmentation pipeline for whole‑slide histopathology images of colorectal cancer, labeling tumor grades 1‑3 and normal mucosa. It employs dense prediction transformers with multiple encoder backbones, overlapping patches, test‑time augmentation, and an adaptive augmentation policy guided by large language models. The approach, combined with soft‑voting ensembles and post‑processing refinements, raises the F1 score from 62.92 to 69.84 on a colorectal cancer grade dataset.

By \"Umit Mert \c{C}a\u{g}lar, Alptekin Temizel
arXiv Computer Vision
Sep 24

CasCVS-Net: A Staged Multi-Task Cascade for Critical View of Safety Assessment

CasCVS‑Net is a staged multi‑task cascade that jointly performs object detection, semantic segmentation, and Critical View of Safety (CVS) assessment for laparoscopic cholecystectomy. The model couples tasks through predicted anatomy—boxes guide segmentation and masks provide region‑level features for CVS classification—allowing CVS assessment to rely solely on model predictions. Trained on the Endoscapes dataset, CasCVS‑Net outperforms state‑of‑the‑art methods, achieving higher mAP and mIoU scores across detection, segmentation, and CVS tasks, especially for rare hepatocystic structures.

By Bock-Zien Toh, Yuanchuan Ren, Tay Aw Yu, Ng Khee Ong, Zhehua Mao, Sophia Bano
arXiv Computer Vision
Sep 22

SurgMotion: A Video-Native Foundation Model for Universal Understanding of Surgical Videos

SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.

By Jinlin Wu, Felix Holm, Chuxi Chen, An Wang, Yaxin Hu, Xiaofan Ye, Zelin Zang, Miao Xu, Lihua Zhou, Huai Liao, Danny T. M. Chan, Ming Feng, Wai S. Poon, Hongliang Ren, Dong Yi, Nassir Navab, Gaofeng Meng, Jiebo Luo, Hongbin Liu, Zhen Lei
arXiv Computer Vision
Sep 16

Lesion-centered 3D mapping of colonoscopy procedures: validation of a hierarchical ensemble pipeline on public benchmark videos

The paper presents a four‑layer hierarchical pipeline that constructs a lesion‑centered spatial record from colonoscopy videos without full‑colon 3D reconstruction. It combines a global topological map, lesion‑level spatio‑temporal tracks, on‑demand local 3D reconstruction, and persistent lesion identity across repeated observations, and evaluates the system on four public videos. The results show successful detection of revisit events, accurate lesion identity merging, and superior geometry accuracy compared to a general‑purpose foundation model.

By Hyunjun Kim, Hyeonwoo Na, Jaewoo Lee
arXiv AI
Jun 17

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.

By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath