arXiv:2606. 17020v1 Announce Type: cross Abstract: Remote sensing vision-language models have advanced Earth observation understanding, but most existing work remains centered on RGB imagery, leaving the complementary information in infrared data underexplored.
By Jiaju Han, Ben Zhang, Xuemeng Sun, Qike Zhang, Yuxian Dong, Chengyin Hu, Fengyu Zhang, Yiwei Wei, Jiujiang Guo
Controllable infrared-visible image fusion aims to integrate complementary thermal and structural information with flexible region-aware modulation, producing fused images that adapt to diverse user requirements and downstream tasks. However, existing methods typically rely on predefined discrete control conditions, leading to a sparse space that fails to support fine-grained modulation demands.
The paper introduces OmniRSCLIP, an end‑to‑end contrastive learning framework that extends the CLIP architecture to handle heterogeneous remote sensing sensors such as SAR, multi‑spectral imaging, and hyperspectral imaging. It achieves this by employing Spectral‑Spatial Basis Decomposition to adapt arbitrary‑channel inputs without losing pretrained visual knowledge, and a spectral‑context‑aware mask‑based contrastive learning scheme to improve fine‑grained image‑text alignment. The authors also build OmniRS5M, a large‑scale image‑text corpus covering multiple sensor modalities, and demonstrate that OmniRSCLIP maintains strong RGB performance while effectively supporting these diverse remote sensing data types.
By Xiangyang Miao, Kelu Yao, Yekai Huang, Xiaogang Xu, Junxiao Xue, Minjun Shen, Chenghui Lv, Shanji Liu, Yaying Chen, Chao Li
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2606. 10819v1 Announce Type: cross Abstract: RS-MLLMs enable natural-language understanding and spatial reasoning over earth observation imagery.
By Miaoxin Cai, Guanqun Wang, Wei Zhang, Guangyao Zhou, Yin Zhuang, Tong Zhang, Hao Wang, He Chen, Jun Li
arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.
By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
arXiv:2605.17949v2 Announce Type: replace
Abstract: Remote sensing vision-language models (RS-VLMs) commonly employ a pretrained vision encoder and a projection module to map image features into the...
By Xiao Yang, Ronghao Fu, Zhiwen Lin, Zhuoran Duan, Lang Sun, Jiaqi Liu, Jiashun Zhu, Jiasen Hu, Xu Na, Bo Yang
The paper introduces DOD-SA, a framework for infrared-visible object detection that uses only single-modality annotations. It employs a Collaborative Teacher-Student Network with a single-modality branch and a dual-modality decoupled branch to transfer knowledge across modalities, and a Progressive and Self‑Tuning Training Strategy to refine pseudo‑labels. A Pseudo Label Assigner is also designed to align labels between modalities during training.
By Hang Jin, Chenqiang Gao, Junjie Guo, Fangcen Liu, Qinyao Chang, Kanghui Tian, Deyu Meng
arXiv:2607. 29445v1 Announce Type: cross Abstract: Infrared vision-language models (IR-VLMs) extend thermal perception to open-vocabulary classification, image captioning, and visual question answering.
By Xiang Chen, Yingying Zhao, Chao Li, Jiaju Han, Ben Zhang, Ang Li, Jiahuan Long, Yiwei Wei, Jiujiang Guo, Chengyin Hu
arXiv:2512.15971v2 Announce Type: replace
Abstract: Multispectral object detection is critical for safety-sensitive applications such as autonomous driving and surveillance, where robust perception u...
By Manuel Nkegoum, Minh-Tan Pham, \'Elisa Fromont, Bruno Avignon, S\'ebastien Lef\`evre
GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.
By Ashutosh Gandhe, Anupam Rawat, Geet Sethi, Kabir Nasiruddin, Madhav Kotecha, Panav Shah, Rakshit Sawarn, Soumitra Nayak
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.