The integration of spatial and spectral information is beneficial to the improvement of change detection performance. However, existing methods cannot efficiently suppress the influences of spatial and spectral differences in unchanged areas.
The core challenge of heterogeneous change detection in remote sensing imagery lies in effectively decoupling genuine land-cover changes from significant modal disparities caused by distinct imaging mechanisms. These intrinsic inconsistencies are prone to introducing pseudo-changes, thereby constraining detection accuracy.
The paper introduces SGFNet, a Semantic‑Guided Fusion Network for classifying multi‑source remote sensing images. It features a Semantic Mixing Convolution Block that generates semantic‑aware kernels based on contextual relationships, and a Frequency Modulated Fusion Block that fuses cross‑modal information in the frequency domain to mitigate spatial misalignment. Experiments on the Augsburg and Houston 2018 datasets show SGFNet consistently outperforms state‑of‑the‑art methods.
By Yuwei Zhao, Chuanzheng Gong, Baogui Huan, Feng Gao, Junyu Dong, Qian Du
The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.
By Yuan Qian, Jie Ma
arXiv:2606. 10329v1 Announce Type: cross Abstract: As one of the most destructive natural disasters, earthquakes have struck many countries around the world in recent years, causing serious economic losses.
By Yunlong Liu, Zekai Zhang
arXiv:2606. 27410v1 Announce Type: cross Abstract: The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points.
By Yelin Wang, Zijia Song, Chuanguang Yang, Miaoyu Wang, Zhulin An, Libo Huang, Yongjun Xu
The paper introduces S$^3$F-Net, a dual‑branch network that fuses spatial and spectral representations for medical image classification. It combines a deep spatial CNN with a shallow spectral encoder, SpectraNet, which uses a learnable SpectralFilter layer to process the full Fourier spectrum efficiently. Evaluated on four medical imaging datasets, S$^3$F-Net consistently outperforms spatial‑only baselines, achieving state‑of‑the‑art accuracy on BRISC2025 and surpassing deeper models on the Chest X‑Ray Pneumonia dataset.
By Md. Saiful Bari Siddiqui, Mohammed Imamul Hassan Bhuiyan
The paper introduces MaSoN, an end-to-end unsupervised remote sensing change detection framework that synthesises diverse changes directly in latent feature space during training. By generating changes based on feature statistics of the target data, MaSoN produces data‑driven variations that align with the target domain and can be applied to new modalities such as SAR and multispectral imagery. The method achieves a 14.1 percentage point improvement in average F1 score across five benchmarks, demonstrating strong generalisation across diverse change types.
By Bla\v{z} Rolih, Matic Fu\v{c}ka, Filip Wolf, Luka \v{C}ehovin Zajc
arXiv:2606. 27018v1 Announce Type: cross Abstract: Remote Sensing Foundation Models (RSFMs) have emerged as a powerful alternative to supervised models for Earth Observation, allowing satellites to autonomously trigger high-resolution captures or adjust tasking parameters upon detecting an anomaly, thereby maximizing the utility of the mission's limited power and computational resources.
By S. Ram\'irez-Gallego
arXiv:2505.15147v3 Announce Type: replace
Abstract: Remote sensing images (RSIs) capture both natural and human-induced changes on the Earth's surface. Semantic segmentation (SS) of RSIs enables the...
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
The paper introduces STAND, a method for remote sensing image change captioning that tackles ambiguities in viewpoint, scale, and prior knowledge. It employs a semantic anchoring constraint to regularize temporal representations, a dual‑granularity disambiguation module that uses global context and frequency‑refocused attention to resolve spatial uncertainties, and a semantic concept anchoring module that leverages language priors during decoding. Experiments demonstrate that STAND outperforms existing approaches and effectively addresses these ambiguities.
By Yanpei Gong, Beichen Zhang, Hao Wang, Xuhang Fu, Zhaobo Qi, Xinyan Liu, Yuanrong Xu, Weigang Zhang
arXiv:2606. 28724v1 Announce Type: cross Abstract: Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing.
By Jinhong Hu, Xiaoping Wang, Shuyin Huang, Guojin Zhong, Kaitai Liu, Kai Lu