arXiv AI

SAGE: Source-Anchored Guidance via Frequency Equalization for Hierarchical RGB-T Alignment and Fusion

The paper introduces SAGE, a unified framework for RGB‑Thermal (RGB‑T) image alignment and fusion that integrates frequency equalization, hierarchical alignment, and subband fusion. SAGE uses invertible joint encoding and source‑specific low‑frequency modulation to generate structural and gain guidance, then performs hierarchical frequency collaborative alignment to estimate global affine geometry and refine high‑frequency details. Guided subband fusion aggregates aligned frequency coefficients under propagated guidance, coordinating complementary low‑ and high‑frequency information to reconstruct the fused image via inverse wavelet transform. Experiments on real‑world and synthetic misaligned RGB‑T datasets show that SAGE achieves competitive performance in both alignment and fusion, validating the effectiveness of source‑anchored guidance for weakly registered RGB‑T images.

arXiv Computer Vision
Sep 24

S2A:Semantic-to-Spatial Alignment for Alignment-Free RGB-T Salient Object Detection

The paper introduces S2A, a semantic-to-spatial alignment framework designed for alignment‑free RGB‑T salient object detection. It employs a global‑guided hierarchical fusion module to refine intra‑modal features, an alignment‑free cross‑modal channel attention module to exchange semantic information, and a spatial deformable cross‑attention module to recover local spatial correspondence. These components collectively reduce misalignment‑induced feature contamination and achieve competitive performance on public benchmarks without additional bells and whistles.

By Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu
arXiv Computer Vision
Sep 24

Diff-RF: Mutually Reinforced Image Registration and Fusion via Degradation-Aware Learning

Diff‑RF is a diffusion‑based framework that jointly performs image registration and fusion while accounting for degradation in multi‑modal images. It first restores modality‑specific degradations within each image, then uses a cross‑modal diffusion module that couples registration and fusion, refining alignment and enhancing complementary information. Experiments on extended datasets show that this coupled approach yields higher registration accuracy and fusion quality under diverse degraded conditions.

By Xunpeng Yi, Zaixi Du, Qinglong Yan, Yibing Zhang, Han Xu, Jiayi Ma
arXiv Computer Vision
Sep 30

End-to-End Self-Supervised RGB-T Tracking without Modality Misleading

ESMTrack is a fully end‑to‑end self‑supervised RGB‑T tracking framework that eliminates the need for costly modality‑aligned bounding boxes or offline pseudo‑label generation. It learns discriminative, temporally consistent representations using a grounding triplet loss on the initial annotated frame and a cross‑frame temporal triplet loss on unlabeled search frames, with reliable samples selected via forward‑backward consistency. A three‑branch architecture (fusion, RGB, thermal) and a modality decoupling mechanism mitigate modality dominance bias, enabling competitive state‑of‑the‑art performance, strong cross‑dataset generalization, and real‑time inference on five RGB‑T benchmarks.

By Shenglan Li, Rui Yao, Kunyang Sun, Hong Jia, Yong Zhou, Javen Qinfeng Shi, Xinyu Zhang
arXiv Computer Vision
Sep 14

RoES: Rotational Equivariant Selective-frequency Fusion for Multimodal Images

RoES is a Rotational Equivariant Selective-frequency fusion network that dynamically separates low- and high-frequency components of infrared-visible images. It uses a trainable rotation-enhanced updater to decouple frequencies, then fuses them with a dual-branch module: a rotation-equivariant Mamba for low-frequency structural dependencies and a polar spectral attention Dual-Fourier block for high-frequency detail refinement. Experiments show RoES outperforms existing methods in fusion quality and downstream object detection, offering a robust multimodal fusion solution.

By Jiabao Wang, Wenjian Liu, Yaoming Cai, Gengyu Zhang, Boyan Zhao, Zijia Zhang, Yao Ding, Xiaobo Liu
arXiv Computer Vision
Sep 4

Residual Optimal Transport-Based Experts Collaboration Towards Modality-Aware Infrared-Visible Object Detection

The paper introduces FlexibleFusion, a method for infrared-visible object detection that adapts to both complete and missing-modality scenarios. It employs a Modality-Aware Experts Collaboration mechanism to selectively fuse cross-modal or intra-modal pathways, and a Residual Self-Paced Entropic Optimal Transport module to align heterogeneous feature distributions without heavy optimization. Experiments demonstrate consistent performance across various modality configurations.

By Yue Zhao, Hua Yu, Yukun Zhao, Yuzhi Zhang, Maoguo Gong, Xin Mei, Zhuping Hu, Yanchi Li, A. K. Qin