Diff‑RF is a diffusion‑based framework that jointly performs image registration and fusion while accounting for degradation in multi‑modal images. It first restores modality‑specific degradations within each image, then uses a cross‑modal diffusion module that couples registration and fusion, refining alignment and enhancing complementary information. Experiments on extended datasets show that this coupled approach yields higher registration accuracy and fusion quality under diverse degraded conditions.
By Xunpeng Yi, Zaixi Du, Qinglong Yan, Yibing Zhang, Han Xu, Jiayi Ma
Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures often exhibit substantially different image intensit...
arXiv:2609.15669v1 Announce Type: cross
Abstract: Multimodal image registration is a key component of many clinical workflows, yet it remains challenging because corresponding anatomical structures o...
By Matteo Barbieri, Giammarco La Barbera, Juan Pablo De La Plata, Sabine Sarnacki, Isabelle Bloch, Pietro Gori
CoMLP introduces a cooperatively-gated MLP module that fuses multimodal medical data—such as imaging modalities and clinical reports—without relying on computationally heavy cross-attention. The module uses regional and dilated MLP interactions to capture both local and global cross-modal dependencies, enabling fine-grained fusion at high spatial resolutions. Experiments on five segmentation benchmarks, covering 2D/3D images and diverse anatomical regions, show consistent improvements over state-of-the-art multi-modal and language-guided methods, highlighting the effectiveness of MLP-based interaction for medical image segmentation.
By Mingyuan Meng, Shuchang Ye, Mingjian Li, Zhenyu Zhao, Jinman Kim, Lei Bi
CEM‑TUDASR is a lightweight, unsupervised Transformer-based super‑resolution framework designed to enhance low‑resolution images from Wireless Capsule Endoscopy (WCE). It uses a domain‑adaptive degradation network to generate realistic WCE‑like low‑resolution images from high‑resolution conventional endoscopy data, enabling effective unpaired learning. The model incorporates Deep Attention Blocks and a Fusion Attention Block to capture both global context and fine local details, achieving superior performance on WCE datasets and demonstrating cross‑domain adaptability to retinal images, all while keeping the parameter count and computational load low.
arXiv:2609.24560v1 Announce Type: new
Abstract: Confocal Laser Endomicroscopy (CLE) provides real-time, cellular-resolution optical biopsy but has a narrow field of view, which image mosaicing can ex...
By Ahmed Aboelela, Johannes Barcsay, Jana Friedhof, Miguel Gon\c{c}alves, Alexander Hann, Katharina Breininger