CEM‑TUDASR is a lightweight, unsupervised Transformer‑based super‑resolution framework designed for Wireless Capsule Endoscopy (WCE) images. It uses a domain‑adaptive degradation network to synthesize realistic low‑resolution WCE images from high‑resolution conventional endoscopy data, enabling unpaired training. The SR generator incorporates Deep Attention Blocks and a Fusion Attention Block to preserve both global context and fine local structures, achieving superior no‑reference quality metrics and improved restoration of mucosal textures, vascular patterns, and anatomical details while remaining computationally efficient.
By Anjali Sarvaiya, Jay Kadel, Kishor Upla, Kiran Raja
UnCapsTSR is an unsupervised transformer-based GAN framework designed to enhance the spatial resolution of low‑resolution wireless capsule endoscopy (WCE) images. It eliminates the need for explicit degradation modeling or paired LR‑HR data by using a Bilateral Total Variation loss to preserve spatial continuity. The authors introduce a new Kvasir Capsule dataset for training, validate generalizability on KID and GIANA datasets, and propose the Endoscopy Quality Metric (EndoQM) as a non‑reference evaluation tool, reporting 40–80% improvement in EndoQM over state‑of‑the‑art unsupervised methods.
By Anjali Sarvaiya, Shubh Kawa, Lalit Agrawal, Jagrit Joshi, Kishor Upla, Kiran Raja
arXiv:2608. 07176v1 Announce Type: cross Abstract: Developing foundation generative models for endoscopy is limited by the gap between natural and clinical images and the computational cost of training large Diffusion Transformers.
By Francisco Caetano, Tim J. M. Jaspers, Haiko Middeljans, Martijn R. Jong, Rixta A. H. van Eijck van Heslinga, Floor Slooter, Albert J. de Groof, Jacques J. Bergman, Peter H. N. De With, Fons van der Sommen
MultiAttenGastro is a plug‑and‑play attention framework that adds parallel 1‑D channel, 2‑D spatial, and 3‑D contextual heads to existing CNN and transformer backbones for gastrointestinal endoscopy classification. Across eight backbones and five public GI datasets, the framework improves performance on large‑gap datasets such as Kvasir‑Capsule but shows no benefit on small‑gap benchmarks like Kvasir‑v2, with mixed results elsewhere. Analysis using Centered Kernel Alignment indicates that the gains are linked to representational redundancy: low inter‑head redundancy under large domain gaps yields consistent improvements, while high redundancy under small gaps leads to losses.
By Sadhana Devarajan, Praveen Kumar Chandaliya, Dhruvin Jashvant Kumar Shah, Kishor Upla, Kiran Raja
Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
arXiv:2608. 15537v1 Announce Type: cross Abstract: Accurate boundary delineation remains a persistent challenge in dermoscopic image segmentation because of blurred lesion margins, heterogeneous textures, and complex background artifacts.
By Wang Jiangtao, Nur Intan Raihana Ruhaiyem, Fu Panpan, Yang Yu, Huang Yan