The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.
By Yuan Qian, Jie Ma
The paper introduces MGRL-RSCC, a multi‑granularity reward reinforcement learning framework for Remote Sensing Change Captioning (RSCC). It uses a CNN with hierarchical self‑attention to extract visual features, a Transformer decoder for visual‑to‑linguistic translation, and a dual‑decoding strategy combined with token‑level supervised learning and self‑critical reinforcement learning. Three reward functions—linguistic fluency, change state consistency, and structural‑semantic relevance—are employed to improve caption quality and reduce exposure bias and conservative generation.
By Futian Wang, Mengqi Wang, Xiao Wang, Wentao Wu, Haowen Wang, Zhicheng Zhao, Jin Tang
arXiv:2605. 15375v2 Announce Type: replace-cross Abstract: Remote sensing change detection (RSCD) localises changes between two images of the same geographic region.
By Bla\v{z} Rolih, Matic Fu\v{c}ka, Filip Wolf, Luka \v{C}ehovin Zajc
arXiv:2606. 00987v1 Announce Type: cross Abstract: Large Vision-Language Models (LVLMs) have shown strong visual understanding and language-guided grounding abilities, yet their capacity for multi-temporal visual reasoning remains underexplored.
By Bingyu Li, Da Zhang, Tao Huo, Zhiyuan Zhao, Junyu Gao, Xuelong Li
arXiv:2608. 06150v1 Announce Type: new Abstract: Earth-surface monitoring requires change detection models capable of recognizing arbitrary semantic categories.
By Zijie Wang, Chen Zhong, Wei He
arXiv:2609.36616v1 Announce Type: new
Abstract: Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past ap...
By Hanwen Lu, Jun He, Mingjia Yang, Hao Wei, Jinhao Huang, Yi Lin, Xiang Zhang
arXiv:2607. 02551v1 Announce Type: cross Abstract: Video multimodal large language models have made strong progress on open-ended video understanding, but they still lack precise local spatiotemporal perception.
By Yankai Yang, Yancheng Long, Bin Wen, Fan Yang, Tingting Gao, Han Li, Shuo Yang
arXiv:2608.28247v1 Announce Type: new
Abstract: Change detection in Earth observation (EO) is critical for monitoring land surface transformations, yet recent research in the field is constrained by...
By Tadej Tomani\v{c}, Alice Baudhuin, Jan Soto\v{s}ek, Jure Brence, Pan\v{c}e Panov, Nikola Simidjievski, Dragi Kocev
arXiv:2608. 01856v2 Announce Type: replace Abstract: Bi-temporal remote-sensing disaster change captioning often needs to identify sparse and spatially localized changes across large pre- and post-event scenes and then translate them into coherent, factual descriptions.
By Dongwei Sun, Bowen Yao, Yujie Zhang, Pei Liu, Jing Yao, Xiangyong Cao
arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.
By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
Historical street-view imagery records urban evolution, but uneven coverage leaves substantial gaps in historical records. Generating plausible past appearances requires restoring changed structures w...
arXiv:2607. 22705v1 Announce Type: cross Abstract: Object-centric learning aims to represent scenes as objects whose properties can be reused in new combinations.
By Anuraag Gadehothur Karnam, Tarunesh Sathish