Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.
arXiv:2608. 03023v1 Announce Type: cross Abstract: Remote sensing semantic segmentation is hindered by costly pixel-level annotations, motivating training-free open-vocabulary methods.
By Changhao Zhao, Haoxiang Li, Yuke Li, Hai Liu, LingLin Zeng
arXiv:2609.22834v1 Announce Type: new
Abstract: Open-vocabulary semantic segmentation (OVS) of remote sensing imagery is a challenging pixel-level task requiring strong generalization and adaptation...
By Changhao Zhao, Linglin Zeng, Hai Liu
The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
Rapid advancements in vision-language models have propelled Referring Remote Sensing Image Segmentation (RRSIS) to the forefront of Earth observation. However, practical deployments suffer severe perf...
OptiSAR-Net++ introduces a new cross‑domain remote sensing visual grounding task (CD‑RSVG) and the first large‑scale benchmark dataset, OptSAR‑RSVG. The framework replaces Transformer decoding with a CLIP‑based contrastive approach, employing a patch‑level Low‑Rank Adaptation Mixture of Experts for efficient cross‑domain feature decoupling and a text‑guided dual‑gate fusion module for improved semantic‑visual alignment. Experiments show state‑of‑the‑art performance on OptSAR‑RSVG and DIOR‑RSVG, with notable gains in localization accuracy and computational efficiency.
By Xiaoyu Tang, Jun Dong, Jintao Cheng, Rui Fan
arXiv:2606.16996v2 Announce Type: replace-cross
Abstract: Segment Anything Model 3 (SAM 3) provides a strong frozen backbone for concept-prompted segmentation, but applying it directly to open-vocabu...
By Tran Dinh Tien, Zhiqiang Shen
arXiv:2505.15147v3 Announce Type: replace
Abstract: Remote sensing images (RSIs) capture both natural and human-induced changes on the Earth's surface. Semantic segmentation (SS) of RSIs enables the...
By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
The paper introduces OmniRSCLIP, an end‑to‑end contrastive learning framework that extends the CLIP architecture to handle heterogeneous remote sensing sensors such as SAR, multi‑spectral imaging, and hyperspectral imaging. It achieves this by employing Spectral‑Spatial Basis Decomposition to adapt arbitrary‑channel inputs without losing pretrained visual knowledge, and a spectral‑context‑aware mask‑based contrastive learning scheme to improve fine‑grained image‑text alignment. The authors also build OmniRS5M, a large‑scale image‑text corpus covering multiple sensor modalities, and demonstrate that OmniRSCLIP maintains strong RGB performance while effectively supporting these diverse remote sensing data types.
By Xiangyang Miao, Kelu Yao, Yekai Huang, Xiaogang Xu, Junxiao Xue, Minjun Shen, Chenghui Lv, Shanji Liu, Yaying Chen, Chao Li
arXiv:2606. 28410v1 Announce Type: cross Abstract: Open-vocabulary semantic segmentation (OVSS) enables text-guided segmentation of unseen objects, breaking fixed-class limitations to achieve open-world understanding.
By Shanwen Wang, Xin Sun, Sirui Wang, Xiao Xiang Zhu
The paper introduces SARTM, a framework that adapts the Segment Anything Model (SAM) for RGB‑thermal (RGB‑T) semantic segmentation. It fine‑tunes SAM with LoRA layers, incorporates language guidance, and employs a Cross‑Modal Knowledge Distillation module to bridge modality gaps. The approach also modifies the segmentation head and adds an auxiliary semantic head, achieving superior performance on MFNET, PST900, and FMB benchmarks.
By Dong Xing, Jinhe Zhang, Hang Yang, Yuqing Wang
arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.
By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")