Reliable Mislabel Detection for Video Capsule Endoscopy Data
arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.
arXiv:2310. 07895v2 Announce Type: replace Abstract: This paper presents a method to efficiently classify the gastroenterologic section of images derived from Video Capsule Endoscopy (VCE) studies by exploring the combination of a Convolutional Neural Network (CNN) for classification with the time-series analysis properties of a Hidden Markov Model (HMM).
arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.
arXiv:2607. 22173v1 Announce Type: cross Abstract: Bowel obstruction is a common and potentially life-threatening gastrointestinal condition.
CapsuleMotion is a lightweight, real‑time visual motion predictor designed for video capsule endoscopy (VCE). It predicts motion between successive frames using on‑device image compression metrics, allowing the capsule to adjust its frame rate dynamically and operate in a low‑power mode before entering the small intestine. Evaluated on the Rhode Island VCE dataset and deployed on an ultra‑low‑power RISC‑V demonstrator, CapsuleMotion reduces energy consumption by up to 20.66% and improves detection of the small intestine entry point.
UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.
arXiv:2511. 01143v2 Announce Type: replace-cross Abstract: Early and accurate segmentation of colorectal polyps is critical for reducing colorectal cancer mortality, which has been extensively explored by academia and industry.
WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.
arXiv:2508. 17728v2 Announce Type: replace-cross Abstract: Cervical cancer remains a significant global health concern and a leading cause of cancer-related deaths among women.
Colorectal cancer (CRC) is the second most deadly and third most common cancer, and the leading cause of death among gastrointestinal cancers. Early diagnosis is crucial for the treatment of this canc...
arXiv:2607. 10357v1 Announce Type: cross Abstract: The automatic detection and classification of cardiovascular disease (CVD) from computed tomography (CT) images plays an important role in clinical practice.
The paper presents an automated segmentation pipeline for whole‑slide histopathology images of colorectal cancer, labeling tumor grades 1‑3 and normal mucosa. It employs dense prediction transformers with multiple encoder backbones, overlapping patches, test‑time augmentation, and an adaptive augmentation policy guided by large language models. The approach, combined with soft‑voting ensembles and post‑processing refinements, raises the F1 score from 62.92 to 69.84 on a colorectal cancer grade dataset.
arXiv:2610.00414v1 Announce Type: new Abstract: Foundation models pretrained on large-scale datasets demonstrate strong transferability to medical imaging tasks. However, understanding how their late...
arXiv:2608. 07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers.