Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.
arXiv:2609.24539v1 Announce Type: new
Abstract: Multi-modal object Re-Identification (ReID) benefits from complementary information across heterogeneous imaging modalities. To further enrich semantic...
By Weixiang Zhou, Yuhao Wang, Xingguo Xu, Weizhen Zhou, Zhixun Su, Jinshan Pan, Cong Wang
Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored. Infrared images provide distinctive cues, including thermal intensity structures, object boundaries, and illumination-invariant scene features, which can enrich visual-language learning beyond conventional RGB observations.
arXiv:2609.38968v1 Announce Type: new
Abstract: Multi-modal image fusion (MMIF) aims to form a single image by integrating shared information, preserving complementary cues, and coordinating cross-mo...
By Zeyu Wang, Mingyu Ge, Haiyu Song, Haoran Duan
The paper introduces IT-TextFusion, an iterative text-guided image fusion framework that uses text-conditioned feature interaction across multiple fusion and refinement stages. It incorporates deep cross-attention, multi-scale cross-gate fusion, and stage-specific text-conditioned modulation to enable degradation-aware global semantic conditioning while preserving complementary visible and infrared information. Experiments on benchmark datasets demonstrate improvements in information-preservation and perceptual-quality metrics, with some metric-dependent trade-offs.
By Siyang Liu, Peiyi Zhou, Tianle Jin, Rongrong Bian, Zheke Jin, Mengze Gao
arXiv:2606. 17020v1 Announce Type: cross Abstract: Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored.
By Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang, Yuxian Dong, Chengyin Hu, Fengyu Zhang, Yiwei Wei, Jiujiang Guo