arXiv:2608. 03826v1 Announce Type: cross Abstract: Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues.
By Jiapeng Li, Yong Li, Junjie Zhou, Fan Zhang, Yu Liu
The paper introduces RSMEB, a unified benchmark for remote‑sensing multimodal retrieval that evaluates both cross‑modal and interleaved retrieval across 21 tasks under a single ranking protocol. It also presents VLM2GeoVec, an instruction‑conditioned single‑encoder model that embeds image, text, bounding‑box, and geo‑coordinate tokens into one sequence and achieves state‑of‑the‑art performance on region‑caption, referring‑expression, and semantic geo‑aware retrieval while remaining competitive on conventional tasks. The authors provide code, checkpoints, and data on GitHub to facilitate reproducibility.
By Emanuel S\'anchez Aimar, Gulnaz Zhambulova, Fahad Shahbaz Khan, Yonghao Xu, Michael Felsberg
arXiv:2512.02697v4 Announce Type: replace
Abstract: Cross-view geo-localization infers a location by retrieving geo-tagged reference images matching a query image. However, the traditional satellite-...
By Zixuan Song, Jing Zhang, Di Wang, Zhiming Luo, Wenbin Liu, Haonan Guo, En Wang, Bo Du, Liangpei Zhang
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2607. 16305v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have achieved strong progress in multimodal understanding.
By Zeyu Xu, Xingzhong Hou, Pengkai Guo, Siling Lin, Xiao Xu, Menghua Zhai, Haoyu Chen, Yunke Zhang, Fei Huang
arXiv:2604.12335v2 Announce Type: replace-cross
Abstract: Training multimodal large language models (MLLMs) for video understanding requires large-scale annotated data spanning diverse tasks such as...
By Tanzila Rahman, Renjie Liao, Leonid Sigal
arXiv:2609.37225v1 Announce Type: cross
Abstract: Multimodal large language models (MLLMs) have shown strong potential for universal multimodal representation learning. However, existing methods eith...
By Zijing Cai, Yuzhe Wang, Jingxian Zhu, Fengbin Zhu, Richang Hong
arXiv:2606. 24997v1 Announce Type: new Abstract: Geographic implicit neural representations (INRs) learn to map any coordinate on Earth to a location embedding, implicitly encoding geospatial data into the weights of a neural network.
By Livia Betti, Sebastian Ricke, Ivica Obadic, Adam J. Stewart, Esther Rolf
Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.
Current instruction-based image retrieval systems are powerful but limited to single-turn interactions, failing to capture the iterative nature of complex, real-world visual searches. To overcome this limitation, we introduce Contextual Composed Image Retrieval (CoCo-IR), a novel task that enables users to progressively refine search results through interactions.
arXiv:2607. 26107v1 Announce Type: cross Abstract: Dense vision-language understanding, including object localization, region recognition, and open-vocabulary semantic segmentation, requires associating language concepts with spatially grounded visual regions.
By Xinran Liu, Shouqian Shi, Yutong Chen, Ge Wang, Xin-Wei Yao, Sheng Zhong
RegRet is a large multimodal model framework that improves region-level retrieval by adding a Region‑Aware Encoder and a multi‑stage training pipeline featuring localized captioning and regional contrastive learning. It also introduces the REGMB benchmark, containing 225k contrastive pairs across four multimodal retrieval tasks. Experiments show RegRet surpasses strong baselines in zero‑shot settings and gains over 20% improvement on REGMB and public benchmarks while maintaining global retrieval performance.
By Xun Liang, Honghui Yang, Weihang Pan, Ruisi Zhao, Boyuan Pan, Yao Hu, Wenxiao Wang, Binbin Lin, Deng Cai