arXiv:2605. 14925v2 Announce Type: replace-cross Abstract: Drone-view geo-localization aims to match a query drone image, often captured under adverse weather conditions (e.
By Yunsong Fang, Tingyu Wang, Zhedong Zheng
arXiv:2607. 23024v1 Announce Type: cross Abstract: High-resolution satellite imagery is the backbone of good land-cover classification, and without that, environmental monitoring, urban planning, and sustainable resource management all fall short.
By Atiq Ur Rehman, Joseph Michael Donovan
arXiv:2608. 19766v1 Announce Type: cross Abstract: Self-supervised pretraining on remote sensing imagery typically treats all samples as equally informative, despite large variability in geographic and visual structure.
By Daniele Rege Cambrin, Francesco Rossi, Mattia Varile
arXiv:2609.38603v1 Announce Type: new
Abstract: While earth observation models have advanced substantially, they still lack interpretability. While concept-bottleneck models provide interpretability...
By Rishabh Mondal, Nipun Batra, Utkarsh Mall
The paper proposes a method called Restrict, Don't Retrain that enhances zero-shot aerial segmentation by using inference-time guidance from a vision‑language model (VLM). It combines a frozen foundation model that labels every pixel with two VLM queries: one to select relevant classes and another to locate small objects missed by the base model. Experiments on four aerial datasets show consistent performance gains at each stage where the base model is competent.
By Teresa DiMeola, Charles Walter, Hong Xiao
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.