Multimodal fusion learning (MFL) has shown great potential in the medical domain, where we are faced with disparate data modalities such as imaging, clinical records, and omics. However, existing MFL strategies face several major challenges.
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.
Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity.
The paper introduces a self‑supervised neural network that unifies single‑frame Fresnel coherent diffraction imaging (CDI) and overlapped ptychography. By using a fixed, pre‑estimated probe and optimizing with a Poisson negative log‑likelihood objective, the method reconstructs object patches from either a single diffraction frame or multiple overlapping measurements, achieving high SSIM scores and a ten‑fold improvement in photon‑dose efficiency. Demonstrations on synthetic patterns and real datasets from APS and LCLS show robust, high‑throughput reconstructions, with a 36× speedup over iterative solvers for a 10,304‑frame workload.
By Oliver Hoidn, Steven Henke, Albert Vong, Aashwin Mishra, Apurva Mehta, Matthew Seaberg
arXiv:2607. 24110v1 Announce Type: cross Abstract: Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception.
By Minchong Chen, Xiaoyun Yuan, Minyu Cao, Jianing Zhang, Jun Zhang, Shuyang Liu, Xiaokang Yang
arXiv:2505.16157v3 Announce Type: replace
Abstract: Transformer-based models have made remarkable progress in image restoration (IR) tasks. However, the quadratic complexity of self-attention in Tran...
By Yuang Ai
Mobile infrared-visible imaging typically pairs a compact infrared sensor with a high-resolution visible camera for complementary perception. While cross-sensor misalignment caused by different optics, viewpoints, fields of view, and exposure timings hinders practical deployment.
arXiv:2609.24494v1 Announce Type: new
Abstract: Accurate monocular depth estimation serves as a core enabler for single camera scene understanding. However, existing self-supervised monocular depth e...
By Xuezhi Xiang, Jiayao Liu, Heqi Xiang, Yuqi Hu, Yiming Chen, Shanjun Zhang
arXiv:2608.06205v2 Announce Type: replace
Abstract: Multispectral object detection combines visible and thermal imagery to improve perception under challenging illumination and environmental conditio...
By Nima Hatami, Karim Faez, Saeed Sharifian, Hamidreza Amindavar
CEM‑TUDASR is a lightweight, unsupervised Transformer-based super‑resolution framework designed to enhance low‑resolution images from Wireless Capsule Endoscopy (WCE). It uses a domain‑adaptive degradation network to generate realistic WCE‑like low‑resolution images from high‑resolution conventional endoscopy data, enabling effective unpaired learning. The model incorporates Deep Attention Blocks and a Fusion Attention Block to capture both global context and fine local details, achieving superior performance on WCE datasets and demonstrating cross‑domain adaptability to retinal images, all while keeping the parameter count and computational load low.
arXiv:2511.18028v2 Announce Type: replace
Abstract: Image super-resolution (SR) is a critical technology for overcoming the inherent hardware limitations of sensors. However, existing approaches main...
By Chenyu Li, Danfeng Hong, Bing Zhang, Zhaojie Pan, Naoto Yokoya, Jocelyn Chanussot
CEM‑TUDASR is a lightweight, unsupervised Transformer‑based super‑resolution framework designed for Wireless Capsule Endoscopy (WCE) images. It uses a domain‑adaptive degradation network to synthesize realistic low‑resolution WCE images from high‑resolution conventional endoscopy data, enabling unpaired training. The SR generator incorporates Deep Attention Blocks and a Fusion Attention Block to preserve both global context and fine local structures, achieving superior no‑reference quality metrics and improved restoration of mucosal textures, vascular patterns, and anatomical details while remaining computationally efficient.
By Anjali Sarvaiya, Jay Kadel, Kishor Upla, Kiran Raja