arXiv Machine Learning By Yanpei Gong, Beichen Zhang, Hao Wang, Xuhang Fu, Zhaobo Qi, Xinyan Liu, Yuanrong Xu, Weigang Zhang

STAND: Semantic Anchoring Constraint with Dual-Granularity Disambiguation for Remote Sensing Image Change Captioning

Read the original on arXiv Machine Learning →

The paper introduces STAND, a method for remote sensing image change captioning that tackles ambiguities in viewpoint, scale, and prior knowledge. It employs a semantic anchoring constraint to regularize temporal representations, a dual‑granularity disambiguation module that uses global context and frequency‑refocused attention to resolve spatial uncertainties, and a semantic concept anchoring module that leverages language priors during decoding. Experiments demonstrate that STAND outperforms existing approaches and effectively addresses these ambiguities.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
2d ago

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.

By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang