arXiv AI By Yuheng Zong, Minghua Wang, Xin Zhao, Zhi-Hui Zhan, Antonio Plaza, Jon Atli Benediktsson

Filling Before Advancing: Capability-Gap-Driven Post-Training for Scenario-Specialized Remote Sensing MLLMs

Read the original on arXiv AI →

arXiv:2607. 22205v2 Announce Type: replace-cross Abstract: Remote sensing multimodal large language models (RS-MLLMs) have improved general aerial-image understanding.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
4d ago

RS-OPSD: Reliable Privileged On-Policy-Self-Distillation for Ultra-High-Resolution Remote Sensing VQA

The paper introduces RS-OPSD, a reliable privileged on-policy self-distillation framework designed for ultra‑high‑resolution remote sensing visual question answering. It leverages a new dataset, GeoEvidence‑6K, and a human‑feedback guided skill refinement process to provide explicit question‑relevant evidence. By incorporating context‑preserving visual privilege and correctness‑aligned distillation, RS‑OPSD achieves state‑of‑the‑art performance on several benchmarks without requiring additional visual search or tool calls at inference time.

By Chengjie Jiang, Yunqi Zhou, Jiafeng Yan, Sihang Zhao, Chun Yuan, Jing Li
arXiv Computer Vision
Sep 16

VPRef: A Cross-Domain Benchmark for Referring Remote Sensing Image Segmentation

The paper introduces VPRef, the first cross‑domain benchmark for Referring Remote Sensing Image Segmentation, containing 46,972 language‑image‑annotation triplets with a three‑tier linguistic hierarchy. It proposes a parameter‑efficient adaptation method based on the Segment Anything Model and Low‑Rank Adaptation, using pseudo‑label self‑training for visual drift and random multi‑granularity prompt mixing for textual drift. Experiments show the approach improves cross‑domain segmentation while altering only 1.08 % of the base model’s parameters, offering a strong baseline for future research.

By Quanwei Liu, Tao Huang, Jiaqi Yang, Wei Xiang
arXiv AI
4d ago

AerialDojo-200K: A Large-Scale Benchmark Suite for Open-World Aerial Object-Goal Search

AerialDojo-200K is a large-scale benchmark suite for open-world aerial object-goal search, featuring 42 simulation scenes across four families and 21 types, including urban, natural, infrastructure, and disaster environments. The dataset contains 205,732 task instances—over 100K semantic-goal and over 100K image-goal tasks—each with a collision-free reference trajectory and multi-view video recordings. A unified evaluation framework splits scenes into 21 in-distribution and 21 out-of-distribution sets, and preliminary tests on multimodal large language models show significant room for improvement in general-purpose aerial agents.

By Tongtong Feng, Xin Wang, Haoran Hou, Ren Wang, Weiran Wang, Shaokai Zhu, Ziqi Jia, Hao Wang, Yu-Wei Zhan, Zongyuan Wu, Jinghao Cui, Wenwu Zhu
arXiv Computer Vision
Sep 22

DeCo: Efficient Decouple-to-Couple Learning for Multi-Task Visual Grounding

DeCo introduces an efficient decouple-to-couple learning framework for multi-task visual grounding, addressing conflicts between localization and segmentation tasks. It first applies Task-aware Semantic Decoupling (TSD) to separate shared visual cues into task-specific features guided by salient words, then uses Hybrid Prior Coupling (HPC) to merge sentence-level semantic priors with mask-derived spatial priors for improved grounding. Experiments across multiple natural and remote sensing datasets show that DeCo achieves state‑of‑the‑art performance while requiring only lightweight trainable parameters on a frozen multimodal encoder.

By Xiaoqiang Lu, Licheng Jiao, Long Sun, Yuting Yang, Xu Liu, Lingling Li, Wenping Ma, Fang Liu