Active Learning for Efficient Annotation of Surgical Videos with Weak Supervision
arXiv:2607. 13237v1 Announce Type: cross Abstract: Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge.
WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.
arXiv:2607. 13237v1 Announce Type: cross Abstract: Precise spatial-temporal annotation of laparoscopic videos is time-consuming and requires expert knowledge.
arXiv:2501.12632v3 Announce Type: replace-cross Abstract: Weakly supervised object localization (WSOL) models can predict both the object class and the spatial regions corresponding to the object, wi...
arXiv:2404. 10034v3 Announce Type: replace-cross Abstract: Weakly Supervised Object Localization (WSOL) allows training deep learning models for classification and localization (LOC) using only global class-level labels.
arXiv:2511. 01143v2 Announce Type: replace-cross Abstract: Early and accurate segmentation of colorectal polyps is critical for reducing colorectal cancer mortality, which has been extensively explored by academia and industry.
arXiv:2608.29759v1 Announce Type: cross Abstract: We present SynCrash, a multi-stage pipeline for zero-shot accident detection, spatial localization, and collision-type classification in fixed-view C...
arXiv:2609.23961v1 Announce Type: new Abstract: Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflectio...
Monocular colonoscopic 3D reconstruction is important for surgical robotic colonoscopy, but remains challenging due to weak texture, specular reflections, limited view overlap, and non-rigid tissue mo...
SurgMotion is a video-native foundation model that replaces pixel-level reconstruction with latent motion prediction for surgical video analysis. It introduces motion-guided masked prediction, spatiotemporal affinity self-distillation, and spatiotemporal feature diversity regularization to focus on semantically meaningful regions and avoid representation collapse. Trained on SurgMotion-15M, the largest surgical video dataset, it outperforms state-of-the-art methods across 17 benchmarks, improving workflow recognition, action triplet recognition, skill assessment, polyp segmentation, and depth estimation.
The paper presents a method for localizing functional surgical landmarks—specifically instrument tips and anchors—in surgical videos without requiring manual pixel-level mask annotations. It leverages vision foundation models, such as SAM 3, to generate dense structural priors through zero‑shot, point‑prompted masks, and refines landmark predictions with a lightweight, coarse‑to‑fine multi‑frame network. Experiments on 7,867 clips from 60 videos show that the approach achieves F1 scores of 72.4% for tip and 58.0% for anchor localization, with ablations confirming the benefits of structural priors and refinement stages.
arXiv:2609.39785v1 Announce Type: new Abstract: The Segment Anything Model (SAM) relies heavily on massive manual annotations, creating a fundamental bottleneck for model scaling. While unsupervised...
arXiv:2609.14466v1 Announce Type: new Abstract: Video segmentation models recognize and track objects over time, but they do not indicate whether each segmented region moves independently of the obse...
DiDA introduces a lightweight video object segmentation framework that leverages Distillation Learning of Deformable Attention. The method uses deformable attention to adapt key and value positions across frames, enabling object representations that are responsive to spatial and temporal changes. Experiments on DAVIS and YouTube‑VOS benchmarks show state‑of‑the‑art performance and efficient memory usage.