Endoscopic Depth Estimation Based on Deep Learning: A Survey
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.
arXiv:2411. 17790v3 Announce Type: replace-cross Abstract: Accurate 3D mapping in endoscopy enables quantitative, holistic lesion characterization within the gastrointestinal (GI) tract, requiring reliable depth and pose estimation.
arXiv:2609.01172v1 Announce Type: new Abstract: Monocular depth estimation has long stood as a fundamental challenge in computer vision, enabling a wide range of applications including 3D reconstruct...
This survey reviews recent advances in surgical video generation, categorizing methods into unconditional, conditional, and world modeling generation. It highlights a shift from creating visually plausible frames to modeling the causal dynamics of surgical scenes, and discusses challenges such as pixel-level fidelity versus clinical plausibility, generalization, physical realism, controllability, and interpretability. The paper also compiles experimental results from public datasets to serve as a quantitative benchmark for the field.
arXiv:2609.02717v1 Announce Type: new Abstract: Large-scale training and refined optimization techniques have greatly improved sparse multi-view 3D reconstruction. Despite their relevance to surgery,...
Effective multi-task learning for surgical scene understanding is fundamentally hindered by annotation granularity mismatch; temporal workflow tasks such as phase recognition, step recognition and anticipation benefit from dense frame-level supervision, whereas pixel-level spatial tasks including instrument segmentation and action recognition are only sparsely annotated on selected keyframes due to prohibitive labeling costs. This supervision imbalance undermines shared representation learning and limits joint optimization across heterogeneous surgical tasks.