DPSF-Net is a dual‑prior spatial‑frequency network designed for real‑world remote sensing image dehazing. It combines hazy RGB images with dark channel prior maps as joint inputs, and incorporates a spatial‑frequency residual interaction block, a prior‑guided feature attention module, and a selective kernel complementary fusion module to reduce colour shift, structural distortion, and large‑scale haze. Experiments show that DPSF-Net achieves state‑of‑the‑art performance on the RRSHID benchmark while maintaining a favorable balance of restoration quality, parameter count, and computational complexity.
By Mei Lu, Shangliang Shao, Shanliang Yao
Infrared-visible image fusion (IVIF) is pivotal for multimodal perception, yet reconciling the inherent information disparity between thermal and textural features remains a fundamental challenge. Existing prior-guided methods often rely on static constraints that induce optimization conflicts or utilize extrinsic semantic priors from large-scale foundation models (e.
arXiv:2608.29220v1 Announce Type: new
Abstract: Multimodal image fusion (MMIF) aims to integrate complementary sensor data into a single representation that preserves intrinsic scene reality while el...
By Haozhen Wei, Chengjun Jiang, Yutong Guo, Xinrui Ju, Xingyuan Li, Xiang Chen, Jinyuan Liu
Multi-modality image fusion (MMIF) enhances scene representation by exploiting complementary cues from different modalities. Adverse weather, however, causes significant image degradation, disrupting feature representation and requiring simultaneous feature restoration and cross-modal complementarity.
RA‑SOD is a new RGB‑Thermal salient object detection framework that explicitly models the reliability of each modality. It introduces a reliability‑conditioned representation, an uncertainty‑guided dual‑stream refinement, and a pixel‑wise modality competition mechanism to adaptively compensate degraded features and suppress unreliable evidence. Experiments on four benchmarks show that RA‑SOD achieves state‑of‑the‑art performance and remains robust under severe modality degradation.
By Hongbo Gao, Zhengyu Li, Xueru Nie, Dihao Zhu, Lijun Zhao, Yunke Wang, Chang Xu
The paper introduces S2A, a semantic-to-spatial alignment framework designed for alignment‑free RGB‑T salient object detection. It employs a global‑guided hierarchical fusion module to refine intra‑modal features, an alignment‑free cross‑modal channel attention module to exchange semantic information, and a spatial deformable cross‑attention module to recover local spatial correspondence. These components collectively reduce misalignment‑induced feature contamination and achieve competitive performance on public benchmarks without additional bells and whistles.
By Qiangqiang Zhou, Yang Luo, Yong Chen, Jiawei Xu
arXiv:2608.29626v1 Announce Type: new
Abstract: Salient object detection in optical remote sensing images (ORSI-SOD) requires dense predictions that preserve object completeness and structural contin...
By Yi Xu, Ruichao Hou, Tongwei Ren, Gangshan Wu
DXPR is a depth‑based cross‑modal place recognition framework that matches monocular camera queries to a LiDAR map using a single vision foundation model backbone. By converting both modalities into a unified depth image representation, DXPR learns modality‑invariant global descriptors without modality‑specific encoders. A geometry‑aware overlap miner refines pairwise metric learning by computing pixel‑level overlap scores, and extensive tests on KITTI and Boreas show strong performance across seasons, weather, and day/night conditions, outperforming prior CMPR baselines.
The paper introduces IT-TextFusion, an iterative text-guided image fusion framework that uses text-conditioned feature interaction across multiple fusion and refinement stages. It incorporates deep cross-attention, multi-scale cross-gate fusion, and stage-specific text-conditioned modulation to enable degradation-aware global semantic conditioning while preserving complementary visible and infrared information. Experiments on benchmark datasets demonstrate improvements in information-preservation and perceptual-quality metrics, with some metric-dependent trade-offs.
By Siyang Liu, Peiyi Zhou, Tianle Jin, Rongrong Bian, Zheke Jin, Mengze Gao
The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.
arXiv:2608.21786v2 Announce Type: replace
Abstract: General image fusion aims to integrate complementary information from multiple source images, but existing methods often rely on task-specific mode...
By Xingxin Xu, Siqi Zhao, Xin Li, Xinjie Yao, Yiming Sun, Pengfei Zhu
The paper introduces SASC-USOD, a framework for underwater salient object detection that learns spatially adaptive coordination between two structural representations: a boundary-sensitive representation using Laplacian filtering and a region-coherent representation via dual-range anisotropic large-kernel aggregation. A spatial coordination module estimates the relative reliability of these representations and adaptively blends them based on image content. Experiments on USOD10K and USOD benchmarks show that SASC-USOD outperforms existing methods, reducing MAE by 4.07% and 23.53% respectively, and its lightweight variant achieves 21 FPS on an NVIDIA Jetson TX2 NX.
By Lin Hong, Chenhui Wang, Linan Deng, Yuning Cui, Yu Zhang, Xin Wang, Bojian Zhang, Xingchen Yang, Fumin Zhang