arXiv Computer Vision

EndoFSA: Endoscopic Few-Shot Image Generation via Rank-Constrained Parameter Adaptation

EndoFSA is a GAN-based model designed for endoscopic few-shot image generation, addressing the scarcity of pathological samples in wireless capsule endoscopy (WCE) data. It adapts a generator pretrained on abundant normal images to abnormal domains by updating only a small set of rank-constrained modulation parameters while keeping the rest of the weights frozen, thereby preserving anatomical priors and preventing mode collapse. The method incorporates perceptual boundary regularization and cluster-wise diversity control, operates without pixel-level annotations, and demonstrates that synthetic abnormal images can match real images in downstream classification performance.

arXiv Computer Vision
Sep 3

UnCapsTSR: An Unsupervised Transformer-based Image Super-Resolution Approach for Capsule Endoscopy Images

UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.

By Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal, Jagrit Joshi, Kishor Upla, Kiran Raja
arXiv AI
Aug 10

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

arXiv:2608. 07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers.

By Francisco Caetano, Tim J. M. Jaspers, Haiko Middeljans, Martijn R. Jong, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Albert J. de Groof, Jacques J. Bergman, Peter H. N. De With, Fons van der Sommen
arXiv Computer Vision
Sep 11

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer‑based super‑resolution framework designed for Wireless Capsule Endoscopy (WCE) images. It uses a domain‑adaptive degradation network to synthesize realistic low‑resolution WCE images from high‑resolution conventional endoscopy data, enabling unpaired training. The SR generator incorporates Deep Attention Blocks and a Fusion Attention Block to preserve both global context and fine local structures, achieving superior no‑reference quality metrics and improved restoration of mucosal textures, vascular patterns, and anatomical details while remaining computationally efficient.

By Anjali Sarvaiya, Jay Kadel, Kishor Upla, Kiran Raja
Hugging Face Trending Papers
Sep 10

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer-based super‑resolution framework designed to enhance low‑resolution images from Wireless Capsule Endoscopy (WCE). It uses a domain‑adaptive degradation network to generate realistic WCE‑like low‑resolution images from high‑resolution conventional endoscopy data, enabling effective unpaired learning. The model incorporates Deep Attention Blocks and a Fusion Attention Block to capture both global context and fine local details, achieving superior performance on WCE datasets and demonstrating cross‑domain adaptability to retinal images, all while keeping the parameter count and computational load low.

arXiv Machine Learning
Jul 8

Reliable Mislabel Detection for Video Capsule Endoscopy Data

arXiv:2602. 06938v2 Announce Type: replace-cross Abstract: The classification performance of deep neural networks relies strongly on access to large, accurately annotated datasets.

By Julia Werner, Julius Oexle, Oliver Bause, Maxime Le Floch, Franz Brinkmann, Hannah Tolle, Jochen Hampe, Oliver Bringmann
arXiv AI
Jun 30

Towards Modality-Agnostic Medical Image Anomaly Detection: A Training-Free Manifold Refinement Approach

arXiv:2604. 19191v2 Announce Type: replace-cross Abstract: Deploying AI-based anomaly detection across diverse clinical imaging settings remains challenging because most existing methods rely on modality-specific architectures, anatomical priors, or extensive retraining, limiting their use as general-purpose screening tools.

By Pritam Kar, Gouri Lakshmi S, Saptarshi Bej
arXiv Computer Vision
Sep 14

Unified CT and MRI Pancreas Segmentation for Label-Efficient Cross-Modality Subregion Transfer

arXiv:2609.13043v1 Announce Type: new Abstract: Robust medical image segmentation across imaging modalities is challenging because of large differences in appearance and intensity distributions. Mode...

By Ziliang Hong, Hongyi Pan, Halil Ertugrul Aktas, Andrea Bejar, Elif Keles, Frank H. Miller, Michael B. Wallace, Rajesh N. Keswani, Gorkem Durak, Ulas Bagci
arXiv Computer Vision
Aug 26

Example-based Robust Abnormality Detection with Minimal Annotations using Exemplar Med-DETR

arXiv:2608.24281v1 Announce Type: new Abstract: Reducing annotation requirements remains a key challenge in developing robust medical object detectors. To address this, Vision-Language (VL) object de...

By Sheethal Bhat, Bogdan Georgescu, Awais Mansoor, Mathias Zinnen, Pranjal Sahu, Florin C. Ghesu, Sasa Grbic, Andreas Maier
arXiv AI
Aug 18

From Generalist to Specialist: A Context-Fusion Framework for Endoscopic Polyp Reporting with a Frozen VLM

arXiv:2608. 15580v1 Announce Type: new Abstract: Reliable endoscopic polyp reporting requires integrating quantitative lesion sizing, standardized Paris classification, and clinically meaningful morphological description within a single record.

By Ruijie Yang, Yan Zhu, Peiyao Fu, Siyuan Li, Te Luo, Zhihua Wang, Quanlin Li, Pinghong Zhou, Xian Yang, Shuo Wang