Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.
arXiv:2608. 03023v1 Announce Type: cross Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods.
By Changhao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
arXiv:2609.22834v1 Announce Type: new
Abstract: Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation...
By Changhao Zhao, Linglin Zeng, Hai Liu
The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe perf...
OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.
By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan