The paper introduces MGRL-RSCC, a multi‑granularity reward reinforcement learning framework for Remote Sensing Change Captioning (RSCC). It uses a CNN with hierarchical self‑attention to extract visual features, a Transformer decoder for visual‑to‑linguistic translation, and a dual‑decoding strategy combined with token‑level supervised learning and self‑critical reinforcement learning. Three reward functions—linguistic fluency, change state consistency, and structural‑semantic relevance—are employed to improve caption quality and reduce exposure bias and conservative generation.
By Futian Wang, Mengqi Wang, Xiao Wang, Wentao Wu, Haowen Wang, Zhicheng Zhao, Jin Tang
The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.
By Yuan Qian, Jie Ma
arXiv:2609.09876v1 Announce Type: new
Abstract: Dense change detection in remote sensing requires vision-language models (VLMs) to compare bi-temporal images and generate accurate pixel-level masks....
By Xiao An, Ruikang Zhang, Chen Zhong, Xuli Shen, Jiaxing Sun, Jiang Wu, Wei He
arXiv:2609.10356v1 Announce Type: new
Abstract: Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban i...
By Benedetta Liberatori, Nermin Samet, Paolo Rota, Matthieu Cord, Elisa Ricci, Andrei Bursuc, Monika Wysocza\'nska
arXiv:2606. 28724v1 Announce Type: cross Abstract: Understanding and localizing subtle changes between paired images is critical for tasks such as surveillance and image editing.
By Jinhong Hu, Xiaoping Wang, Shuyin Huang, Guojin Zhong, Kaitai Liu, Kai Lu
The paper introduces Seeing Before Synthesizing (SBS), a weakly-supervised dense video captioning framework that uses a vision‑language model to generate frame‑level narratives for gaps between events and detect transitions based on semantic changes. SBS refines temporal masks by aligning transition points with vision‑language cues, rather than relying on rigidly placed synthetic captions. Experiments on ActivityNet Captions and YouCook2 show that SBS achieves state‑of‑the‑art results in both captioning and localization tasks.
By Ye-Chan Kim, Seunghee Choi, SeungJu Cha, Si-Woo Kim, Hwiseon Kim, Hyungee Kim, Dong-Jin Kim
Long-term change understanding from images of the same place revisited over time is a challenging task with applications in map maintenance and urban infrastructure monitoring. Prior work addresses it...
Visual Information-Guided Parallel Decoding for Diffusion Multimodal Large Language Models introduces the VIG‑Sampler, a method that prioritizes tokens for decoding based on their attention to image tokens and penalizes redundancy in image‑attention distributions. The approach aims to improve the quality of multimodal generation by selecting more informative tokens during diffusion decoding. Experiments on seven captioning and VQA benchmarks with three open‑source dMLLMs show that VIG‑Sampler outperforms the Info‑Gain Sampler by an average of 19.3 CIDEr points and achieves better COCO Caption results using only half as many decoding steps.
By Insu Lee, Wooje Park, Wonseok Shin, Jinwoo Son, Byonghyo Shim
BiMoGen introduces a unified masked discrete diffusion framework for bidirectional motion‑text generation, addressing the limitations of autoregressive models in capturing bidirectional dependencies between language and motion. The approach employs a two‑stage training strategy—decoupled uni‑ and cross‑modal pretraining followed by supervised fine‑tuning—to establish robust cross‑modal correspondence, and incorporates Generation‑Aware Self‑Correction to mitigate error propagation during inference. Experiments on HumanML3D and KIT‑ML show competitive performance on both text‑to‑motion and motion‑to‑text tasks, demonstrating the effectiveness of the proposed training and correction mechanisms.
arXiv:2606. 27410v1 Announce Type: cross Abstract: The primary goal of Remote Sensing Image Change Captioning (RSICC) is to automatically generate descriptions of changes between remote sensing images captured at different time points.
By Yelin Wang, Zijia Song, Chuanguang Yang, Miaoyu Wang, Zhulin An, Libo Huang, Yongjun Xu
Improving video captioning quality typically demands retraining large vision-language models, an expensive and often impractical requirement. Existing training-free alternatives instead ground captions in detected objects to curb hallucination, but apply only a single, fixed correction pass without prioritizing which objects matter most, leaving semantically significant content omitted.
arXiv:2604. 20623v2 Announce Type: replace-cross Abstract: Traditional change detection identifies where changes occur, but does not explain what changed in natural language.
By Roie Kazoom, Yotam Gigi, George Leifman, Tomer Shekel, Genady Beryozkin