arXiv Machine Learning By Stefan Maria Ailuro (INSAIT, Sofia University "St. Kliment Ohridski"), Mario Markov (INSAIT, Sofia University "St. Kliment Ohridski"), Mohammad Mahdi (INSAIT, Sofia University "St. Kliment Ohridski"), Luc Van Gool (INSAIT, Sofia University "St. Kliment Ohridski"), Danda Pani Paudel (INSAIT, Sofia University "St. Kliment Ohridski")

More with Less: a Large Scale Remote Sensing VLM with a Simple Recipe

Read the original on arXiv Machine Learning →

arXiv:2607. 15942v1 Announce Type: cross Abstract: Remote sensing vision-language models are increasingly expected to support open-ended reasoning over Earth Observation data and a variety of tasks.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Aug 14

LongEarth-R1: Benchmarking and Aligning Vision-Language Models for Long-Horizon Earth Observation Reasoning

arXiv:2608. 13344v1 Announce Type: new Abstract: Long-horizon Earth observation reasoning requires models to organize multi-stage geographic evolution, localize spatial changes, detect temporal anomalies, and infer future from extended image sequences.

By Yupan Ding, Jing Xiao, Zhenyuan Zhang, Chaofeng Chen, Liang Liao, Gui-Song Xia, Mi Wang
arXiv Computer Vision
Sep 25

GeoNLI - A Natural Language Interpreter for Satellite Imagery

GeoNLI introduces a unified, modular pipeline that combines advanced SAM variants with multimodal large language models to perform satellite image captioning, visual question answering (VQA), and visual grounding. The EarthMind model achieves strong results on captioning and VQA, while multiple RemoteSAM-SAM and DiffuSAM pipelines are used for grounding, ultimately employing a majority‑voting ensemble across several models. The system reports 82% captioning accuracy, 83.32% VQA accuracy, and 64.94% grounding accuracy, demonstrating improved consistency over task‑specific approaches.

By Ashutosh Gandhe, Anupam Rawat, Geet Sethi, Kabir Nasiruddin, Madhav Kotecha, Panav Shah, Rakshit Sawarn, Soumitra Nayak
arXiv Machine Learning
Sep 23

MGRL-RSCC: Multi-Granularity Reward Reinforcement Learning for Fine-Grained Remote Sensing Change Captioning

The paper introduces MGRL-RSCC, a multi‑granularity reward reinforcement learning framework for Remote Sensing Change Captioning (RSCC). It uses a CNN with hierarchical self‑attention to extract visual features, a Transformer decoder for visual‑to‑linguistic translation, and a dual‑decoding strategy combined with token‑level supervised learning and self‑critical reinforcement learning. Three reward functions—linguistic fluency, change state consistency, and structural‑semantic relevance—are employed to improve caption quality and reduce exposure bias and conservative generation.

By Futian Wang, Mengqi Wang, Xiao Wang, Wentao Wu, Haowen Wang, Zhicheng Zhao, Jin Tang
arXiv Computer Vision
Sep 24

VLM2GeoVec: Toward Universal Multimodal Embeddings for Remote Sensing

The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.

By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv Computer Vision
Sep 22

DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding

DeCo introduces an efficient decouple-to-couple learning framework for multi-task visual grounding, addressing conflicts between localization and segmentation tasks. It first applies Task-aware Semantic Decoupling (TSD) to separate shared visual cues into task-specific features guided by salient words, then uses Hybrid Prior Coupling (HPC) to merge sentence-level semantic priors with mask-derived spatial priors for improved grounding. Experiments across multiple natural and remote sensing datasets show that DeCo achieves state‑of‑the‑art performance while requiring only lightweight trainable parameters on a frozen multimodal encoder.

By Xiaoqiang Lu, Licheng Jiao, Long Sun, Yuting Yang, Xu Liu, Lingling Li, Wenping Ma, Fang Liu