arXiv Computer Vision

GeoNLI - A Natural Language Interpreter for Satellite Imagery

GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.

arXiv Computer Vision
Sep 24

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.

By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv Machine Learning
Jul 20

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
arXiv Computer Vision
Sep 24

A Unified Framework and Dataset for Oriented Object Visual Grounding in Remote Sensing

The paper introduces O$^2$-VG, a unified framework for oriented object visual grounding in remote sensing images, comprising three complementary models: O$^2$-VG-Trans, a cross‑modality transformer; O$^2$-VG-Uni, which predicts universal oriented proposals; and O$^2$-VG-VLM, an autoregressive vision‑language model that generates oriented bounding boxes. It also presents DIOR‑R‑SVG, a new dataset containing image, expression, and oriented box triplets for training and evaluation. The framework demonstrates superior performance across multiple benchmarks and is supported by publicly available code.

By Zeyu Ding, Yong Zhou, Jiaqi Zhao, Wen-Liang Du, Xixi Li, Hancheng Zhu, Rui Yao, Abdulmotaleb El Saddik
arXiv Machine Learning
Aug 4

OSMDA: OpenStreetMap-based Domain Adaptation for Remote Sensing VLMs

arXiv:2603. 11804v3 Announce Type: replace-cross Abstract: Vision-Language Models (VLMs) adapted to remote sensing rely heavily on domain-specific image-text supervision, yet high-quality annotations for satellite and aerial imagery remain scarce and expensive to produce.

By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Delyan Boychev (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")
arXiv AI
Jun 9

AgroOmni: A Large-Scale Multi-view Agricultural Dataset for Cross-Scale Multimodal Reasoning

arXiv:2603. 14342v2 Announce Type: replace-cross Abstract: Modern agricultural data is sourced from diverse platforms and spans multiple spatial scales, ranging from ground-level close-up photography to Unmanned Aerial Vehicle (UAV) aerial observation and satellite remote sensing imagery.

By Jiarui Zhang, Junqi Hu, Zurong Mai, Yang Liu, Yuhang Chen, Shuohong Lou, Henglian Huang, Hong Cheng, Lingyuan Zhao, Jianxi Huang, Yutong Lu, Haohuan Fu, Juepeng Zheng
arXiv Computation and Language
Sep 17

PANORAMA: Panoptic Grounded Captioning via Mask Proposal Selection

PANORAMA introduces a new panoptic grounded captioning framework that jointly generates detailed image captions and associates each phrase with precise pixel-level masks. The authors create PanoCaps, a human‑annotated benchmark with dense captions and near‑complete pixel coverage, and propose a phrase‑mask matching protocol with a generalized Panoptic Quality metric. PANORAMA conditions a pretrained segmenter on contextualized phrase representations, learns to select appropriate masks, and achieves state‑of‑the‑art grounding performance on PanoCaps and other pixel‑level tasks.

By Sara Pieri, Evangelos Kazakos, Shizhe Chen, Josef Sivic, Cordelia Schmid