arXiv Machine Learning By Julia Werner, Christoph Gerum, Moritz Reiber, J\"org Nick, Oliver Bringmann

Precise localization within the GI tract by combining classification of CNNs and time-series analysis of HMMs

Read the original on arXiv Machine Learning →

arXiv:2310. 07895v2 Announce Type: replace Abstract: This paper presents a method to efficiently classify the gastroenterologic section of images derived from Video Capsule Endoscopy (VCE) studies by exploring the combination of a Convolutional Neural Network (CNN) for classification with the time-series analysis properties of a Hidden Markov Model (HMM).

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 8

Reliable Mislabel Detection for Video Capsule Endoscopy Data

arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.

By Julia Werner, Julius Oexle, Oliver Bause, Maxime Le Floch, Franz Brinkmann, Hannah Tolle, Jochen Hampe, Oliver Bringmann
arXiv Computer Vision
Sep 15

CapsuleMotion: A Lightweight Real-Time Visual Motion Predictor for Capsule Endoscopy

CapsuleMotion is a lightweight, real‑time visual motion predictor designed for video capsule endoscopy (VCE). It predicts motion between successive frames using on‑device image compression metrics, allowing the capsule to adjust its frame rate dynamically and operate in a low‑power mode before entering the small intestine. Evaluated on the Rhode Island VCE dataset and deployed on an ultra‑low‑power RISC‑V demonstrator, CapsuleMotion reduces energy consumption by up to 20.66% and improves detection of the small intestine entry point.

By Oliver Bause, Julia Werner, Oliver Bringmann
arXiv Computer Vision
Sep 3

UnCapsTSR: An Unsupervised Transformer-based Image Super-Resolution Approach for Capsule Endoscopy Images

UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.

By Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal, Jagrit Joshi, Kishor Upla, Kiran Raja
arXiv AI
Sep 10

WSPolypNet: Weakly Supervised Polyp Localization in Colonoscopy Videos

WSPolypNet is a weakly supervised framework that localizes polyps in colonoscopy videos using only video-level labels, avoiding costly frame-level annotations. It employs a 3D CNN to generate class activation maps, enhances them with a multi-view strategy, and refines the results with MedSAM2 segmentation. The method achieves higher CorLoc scores—up to 47.80% at IoU 0.3—and a recall of 94.51%, especially improving detection of small polyps.

By Giseong Hwang, Minjae Jo, Yeonghyeon Park, Kyeonghun Kim, Seoyeon Han, Donghoon Han, Haneul Kim, Yului Jeong, Insung Hwang, Pa Hong, Ken Ying-Kai Liao, Nam-Joon Kim