arXiv AI By Francisco Caetano, Tim J. M. Jaspers, Haiko Middeljans, Martijn R. Jong, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Albert J. de Groof, Jacques J. Bergman, Peter H. N. De With, Fons van der Sommen

Representation-driven Endoscopic Visual Embedding Alignment for Latent Generation

Read the original on arXiv AI →

arXiv:2608. 07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 17

Geometry-Consistent Endoscopic Representations for Image-Guided Navigation via Structured Foundation Model Adaptation

arXiv:2606. 17340v1 Announce Type: cross Abstract: Accurate vision-based navigation in monocular endoscopy is difficult due to limited depth cues, weak tissue texture, non-rigid deformation, and substantial appearance variation across domains, all of which complicate pose estimation, depth prediction, and image-to-anatomy alignment.

By Hongchao Shu, Roger D. Soberanis-Mukul, Hao Ding, Morgan Ringel, Mali Shen, Saif Iftekar Sayed, Hedyeh Rafii-Tari, Mathias Unberath
arXiv Computer Vision
Sep 25

EndoFSA: Endoscopic Few-Shot Image Generation via Rank-Constrained Parameter Adaptation

EndoFSA is a GAN-based model designed for endoscopic few-shot image generation, addressing the scarcity of pathological samples in wireless capsule endoscopy (WCE) data. It adapts a generator pretrained on abundant normal images to abnormal domains by updating only a small set of rank-constrained modulation parameters while keeping the rest of the weights frozen, thereby preserving anatomical priors and preventing mode collapse. The method incorporates perceptual boundary regularization and cluster-wise diversity control, operates without pixel-level annotations, and demonstrates that synthetic abnormal images can match real images in downstream classification performance.

By Panagiota Gatoula, Grigoris Karypidis, Dimitris K. Iakovidis
arXiv Computer Vision
Sep 11

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer‑based super‑resolution framework designed for Wireless Capsule Endoscopy (WCE) images. It uses a domain‑adaptive degradation network to synthesize realistic low‑resolution WCE images from high‑resolution conventional endoscopy data, enabling unpaired training. The SR generator incorporates Deep Attention Blocks and a Fusion Attention Block to preserve both global context and fine local structures, achieving superior no‑reference quality metrics and improved restoration of mucosal textures, vascular patterns, and anatomical details while remaining computationally efficient.

By Anjali Sarvaiya, Jay Kadel, Kishor Upla, Kiran Raja
Hugging Face Trending Papers
Sep 10

CEM-TUDASR: Computationally efficient multi-modality transformer based unsupervised domain adaptive super-resolution approach

CEM‑TUDASR is a lightweight, unsupervised Transformer-based super‑resolution framework designed to enhance low‑resolution images from Wireless Capsule Endoscopy (WCE). It uses a domain‑adaptive degradation network to generate realistic WCE‑like low‑resolution images from high‑resolution conventional endoscopy data, enabling effective unpaired learning. The model incorporates Deep Attention Blocks and a Fusion Attention Block to capture both global context and fine local details, achieving superior performance on WCE datasets and demonstrating cross‑domain adaptability to retinal images, all while keeping the parameter count and computational load low.