arXiv Computer Vision

CrossTimeEdit: A Decade-Spanning Cross-View Dataset and Reward-Guided Editing for Historical Street-View Generation

arXiv Machine Learning
Sep 23

MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning

The paper introduces MGRL-RSCC, a multi‑granularity reward reinforcement learning framework for Remote Sensing Change Captioning (RSCC). It uses a CNN with hierarchical self‑attention to extract visual features, a Transformer decoder for visual‑to‑linguistic translation, and a dual‑decoding strategy combined with token‑level supervised learning and self‑critical reinforcement learning. Three reward functions—linguistic fluency, change state consistency, and structural‑semantic relevance—are employed to improve caption quality and reduce exposure bias and conservative generation.

By Futian Wang, Mengqi Wang, Xiao Wang, Wentao Wu, Haowen Wang, Zhicheng Zhao, Jin Tang
arXiv Computer Vision
Sep 24

From Change Captions to Change Detection: Semantic-Appearance Agreement Framework for Remote Sensing Change Detection

The paper proposes a weakly supervised remote sensing change detection method that uses change captions as the sole supervision signal, eliminating the need for pixel‑level change masks. It introduces a caption‑driven generation pipeline to create bi‑temporal image pairs with controlled changes and a Semantic‑Appearance Agreement Framework (SAAF) that fuses caption‑grounded semantic responses with RGB differences for accurate change localization. Experiments on the Flair‑RSGen and WHU‑CDC datasets demonstrate that SAAF outperforms existing limited‑supervision baselines in macro‑averaged IoU and F1 metrics.

By Yuan Qian, Jie Ma
Hugging Face Trending Papers
Jun 1

MT-EditFlow: Reinforcement Learning for Multi-Turn Image Editing with Flow Matching

Recent breakthroughs in instruction-based image editing have captured significant attention, as models are now capable of handling real-world editing demands with the practicality required by everyday users. However, editing models trained primarily for single-turn edits often break down in multi-turn editing--the natural interactive setting where a user iteratively refines an image based on the model's own previous outputs.

arXiv Computer Vision
Sep 7

AdaptVPR: Route-Aware Hard Positive Generation for Robust Visual Place Recognition

AdaptVPR introduces a route-aware generative augmentation framework that creates hard positive examples for Visual Place Recognition (VPR) training. It uses a vision‑language model to assess scene editability, a rule‑based scheduler to select generation routes, and a VPR‑oriented verification scheme to ensure geometric consistency and appearance diversity. The resulting AdaptCities dataset contains 160K verified synthetic hard positives, leading to consistent performance gains across VPR baselines, including up to 9.2% improvement in R@1 under challenging domain shifts.

By Shunpeng Chen, Jingyi Zhang, Changwei Wang, Shengpeng Xu, Yukun Song, Xingtian Pei, Jinzhou Lin, Li Guo, Shibiao Xu
arXiv AI
Jul 20

GeoChrono: Benchmarking and Rethinking Long-Term Temporal Understanding in Remote Sensing

arXiv:2607. 15768v1 Announce Type: cross Abstract: Remote sensing offers an unparalleled vantage point for observing the Earth's long-term surface evolution, yet it demands that a model not only perceive land cover at isolated moments, but also track changes, memorize evolution histories, and reason across time and space.

By Yujie Li, Jiancheng Pan, Zhiwei Wei, Jiuniu Wang, Mugen Peng, Wenjia Xu