arXiv AI By Ziyuan Liu, Ruifei Zhu, Ouqiao Ma, Yuantao Gu

JL1-CC&QA: Extending the JL1-CD Benchmark with Change Captioning and Question Answering

Read the original on arXiv AI →

arXiv:2606. 31745v1 Announce Type: cross Abstract: Remote sensing change detection (CD) traditionally focuses on pixel-level binary segmentation, which identifies where changes occur but neither what nor why.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 24

From Change Captions to Change Detection: Semantic-Appearance Agreement Framework for Remote Sensing Change Detection

The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.

By Yuan Qian, Jie Ma
arXiv Machine Learning
Sep 23

MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning

The paper introduces MGRL-RSCC, a multi‑granularity reward reinforcement learning framework for Remote Sensing Change Captioning (RSCC). It uses a CNN with hierarchical self‑attention to extract visual features, a Transformer decoder for visual‑to‑linguistic translation, and a dual‑decoding strategy combined with token‑level supervised learning and self‑critical reinforcement learning. Three reward functions—linguistic fluency, change state consistency, and structural‑semantic relevance—are employed to improve caption quality and reduce exposure bias and conservative generation.

By Futian Wang, Mengqi Wang, Xiao Wang, Wentao Wu, Haowen Wang, Zhicheng Zhao, Jin Tang
arXiv Computer Vision
Sep 25

GeoNLI - A Natural Language Interpreter for Satellite Imagery

GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.

By Ashutosh Gandhe, Anupam Rawat, Geet Sethi, Kabir Nasiruddin, Madhav Kotecha, Panav Shah, Rakshit Sawarn, Soumitra Nayak