arXiv AI

Multimodal and Multiscale Spatial-Temporal Semantic Search and Recommendation with AI Foundation Models

arXiv:2606. 28369v1 Announce Type: cross Abstract: Semantic search and recommendation of similar documents, such as news and reports about unusual environmental events (e.

Hugging Face Trending Papers
Aug 4

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes.

arXiv Computation and Language
Aug 25

DamageScope: Vision-Language Retrieval at Scale for Disaster Damage Assessment from Satellite Imagery

DamageScope is a retrieval‑augmented framework that combines satellite imagery, Vision‑Language Models (VLMs), and Large Language Models (LLMs) to automate property damage assessment after natural disasters. It uses a Retrieval‑Augmented Generation (RAG) architecture to extract structured visual representations from satellite images, enabling interactive natural language queries. The system introduces a multi‑vector embedding‑based clustering algorithm that improves scalability and reduces indexing time by up to 14×, and a dual‑store data architecture that cuts LLM API calls, lowering operational cost and response latency by roughly 3×.

By Ravi K. Rajendran, Biplob Debnath, Murugan Sankaradas, Srimat T. Chakradhar
arXiv AI
2d ago

Decoding the Disaster: Multi-Task Geospatial Reasoning with Vision-Language Models and Crowdsourced Imagery for Disaster Mapping

The paper introduces GRDisaster, a multi-task geospatial reasoning framework that leverages vision‑language models to interpret, geolocalize, and assess damage in crowdsourced disaster imagery. It builds on a new benchmark dataset of 26,340 images from PhotoMappers, linking volunteer geographic information, street‑view imagery, and remote sensing data across multiple disaster events from 2018 to 2024. GRDisaster combines deterministic and probabilistic cross‑view geolocalization with multi‑view fusion, and introduces spatial reasoning indicators to validate cross‑view matches and quantify disaster severity using expert‑verified annotations.

By Wenping Yin, Fabian Desuer, Ziqi Liu, Naixia Mou, Weijia Li, Pedram Ghamisi, Xiao Xiang Zhu, Hao Li
arXiv Computer Vision
Sep 17

Multi-View Mixture-of-Experts with Vision-Language Reranking for Cross-View Object Geo-Localization

The paper introduces MVLGeo, a unified framework for cross-view object geo-localization that combines multiple viewpoints into a single model. It employs Vision‑Language Reranking to use contextual text from the query view, a multi‑view Mixture‑of‑Experts architecture to share knowledge and reduce redundancy, and an adaptive elliptical prior for positional encoding. Experiments on CVOGL benchmarks show that MVLGeo achieves state‑of‑the‑art performance and robustness to input degradation.

By Xuyu Fan, Qi Ming, Zhu Han, Liuqian Wang, Siyuan Cao, Xiaohan Zhang, Xudong Zhao, Mingjing Zhao, Yuhan Zhang
arXiv Computer Vision
Aug 21

ID-VTG: Image-Disambiguated Video Temporal Grounding

arXiv:2608. 20127v1 Announce Type: new Abstract: Video Temporal Grounding (VTG) faces significant challenges when natural language queries must distinguish between multiple events involving visually similar entities, particularly when relying on fine-grained visual attributes that are difficult to describe accurately in words alone.

By Minghang Zheng, Jingli Wei, Hongyi Yang, Yang Liu
arXiv Computer Vision
Sep 3

MARS: What Retrieval Signals Are Hidden in Multimodal Large Language Models for Text-Video Retrieval?

MARS introduces a multi‑layer, multi‑slot embedding framework for text‑video retrieval that constructs adaptive representation slots by combining hidden states from different decoder layers. By comparing corresponding text and video slots and aggregating their similarities, MARS captures fine‑grained cues that single‑token embeddings miss. A hard‑negative‑aware slot specialization objective further encourages slots to focus on discriminative matching cues, leading to state‑of‑the‑art results on four benchmarks.

By Uicheol Jung, Juyoung Hong, Geuntaek Lim, Yukyung Choi