arXiv Computer Vision

Diff-RF: Mutually Reinforced Image Registration and Fusion via Degradation-Aware Learning

Diff‑RF is a diffusion‑based framework that jointly performs image registration and fusion while accounting for degradation in multi‑modal images. It first restores modality‑specific degradations within each image, then uses a cross‑modal diffusion module that couples registration and fusion, refining alignment and enhancing complementary information. Experiments on extended datasets show that this coupled approach yields higher registration accuracy and fusion quality under diverse degraded conditions.

Hugging Face Trending Papers
Jul 29

Registration-Grounded Spectral Fusion for Unregistered WLI/NBI Endoscopic Lesion Segmentation

White-light imaging (WLI) and narrow-band imaging (NBI) provide complementary views of endoscopic lesions, but their paired observations are often spatially misaligned due to viewpoint changes, tissue deformation, and sequential handheld acquisition. This makes direct WLI/NBI fusion prone to mixing non-corresponding regions and may even degrade segmentation around lesion boundaries.

Hugging Face Trending Papers
Aug 13

P2Fusion: Prompt-based Progressive Infrared-Visible Image Fusion via Dual-Prior Distillation

Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.

arXiv Computer Vision
Sep 7

CoMLP: Cooperatively-Gated MLPs for Fine-Grained Cross-Modal Information Fusion in Medical Image Segmentation

CoMLP introduces a cooperatively-gated MLP module that fuses multimodal medical data—such as imaging modalities and clinical reports—without relying on computationally heavy cross-attention. The module uses regional and dilated MLP interactions to capture both local and global cross-modal dependencies, enabling fine-grained fusion at high spatial resolutions. Experiments on five segmentation benchmarks, covering 2D/3D images and diverse anatomical regions, show consistent improvements over state-of-the-art multi-modal and language-guided methods, highlighting the effectiveness of MLP-based interaction for medical image segmentation.

By Mingyuan Meng, Shuchang Ye, Mingjian Li, Zhenyu Zhao, Jinman Kim, Lei Bi
arXiv Computer Vision
Aug 25

HP-UniIF: Hierarchical Prompt Learning for Unified Image Fusion

HP-UniIF is a unified vision framework that uses diffusion priors and a depth‑wise hierarchical conditional modulation strategy to support heterogeneous image fusion, visual restoration, and downstream perception tasks. The framework introduces task prompt modulation at bottleneck layers, a degradation prompt router at shallow layers, and an application prompt bank at decoding stages to decouple and adapt to different objectives. Experiments across multiple fusion tasks, degradations, and downstream applications show that HP‑UniIF achieves superior performance while maintaining visually faithful results and task‑relevant semantics.

By Xingxin Xu, Siqi Zhao, Xin Li, Xinjie Yao, Yiming Sun, Pengfei Zhu