arXiv AI

Spatial Representation Learning Beyond Pixels: Unifying Raster Data and Vector Semantics for Human-Centric Geospatial Foundation Models

arXiv:2606. 02374v1 Announce Type: new Abstract: Earth Observation (EO) has fundamentally transformed the monitoring of environmental processes and human activities up to planetary scale.

arXiv AI
Aug 19

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.

By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv Machine Learning
Aug 19

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

MoRAX is a lightweight framework that augments geospatial foundation model embeddings with functional structure derived from human mobility data. By incorporating mobility flows, MoRAX preserves the coverage and consistency of existing geospatial models while adding information about functional connectivity among urban regions, enabling zero‑shot deployment in unseen cities. Experiments across four cities in two countries show that the MoRAX teacher model outperforms baseline geospatial models on eight socioeconomic and environmental prediction tasks, and the student model—without direct mobility input—approaches the teacher’s performance.

By Ya Wen, Jixuan Cai, Yulun Zhou, Alec Kirkley
Hugging Face Trending Papers
Aug 11

GeoSeg-OV: Bridging Geospatial Gaps with Structural Guidance for Open-Vocabulary Remote Sensing Segmentation

Open-vocabulary remote sensing segmentation has recently emerged as a promising paradigm that enables pixel-level recognition of arbitrary categories specified by natural language, including classes unseen during training. However, geospatial domain shifts caused by heterogeneous regions, spatial resolutions, and acquisition platforms weaken visual-text matching and limit cross-dataset generalization.

arXiv Machine Learning
Sep 1

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

BEACON is a tri‑modal contrastive learning framework that enriches AlphaEarth embeddings by aligning physical representations from Earth‑observation imagery with semantic POI text and human behavioral POI visitation data, while keeping the deployed model image‑only. In a Houston case study, BEACON outperformed six baselines on nine downstream tasks, achieving up to 43% higher R² for obesity prevalence, 34% for poor mental health, and 22% for median household income under a linear probe.

By Hao Tian, Heng Cai, Yifan Yang
Hugging Face Trending Papers
Aug 4

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes.

arXiv AI
Jun 19

TerraMind: Large-Scale Generative Multimodality for Earth Observation

arXiv:2504. 11171v5 Announce Type: replace-cross Abstract: We present TerraMind, the first any-to-any generative, multimodal foundation model for Earth observation (EO).

By Johannes Jakubik, Felix Yang, Benedikt Blumenstiel, Erik Scheurer, Rocco Sedona, Stefano Maurogiovanni, Jente Bosmans, Nikolaos Dionelis, Valerio Marsocci, Niklas Kopp, Rahul Ramachandran, Paolo Fraccaro, Thomas Brunschwiler, Gabriele Cavallaro, Juan Bernabe-Moreno, Nicolas Long\'ep\'e
arXiv Machine Learning
Sep 10

Observing Health Outcomes Using Remote Sensing Imagery and Geo-Context Guided Visual Transformer

The paper introduces the Geo-Context Guided Visual Transformer, a model that augments remote sensing image analysis with geospatial embeddings and an asymmetric attention module. By converting heterogeneous geospatial variables into patch-aligned representations and assigning geospatial roles to attention heads, the approach improves disease prevalence prediction over existing vision-language and graph-based baselines. Ablation and visualization studies demonstrate its effectiveness and interpretability for health-related remote sensing tasks, especially when comprehensive geospatial data are scarce.

By Yu Li, Guilherme N. DeSouza, Praveen Rao, Chi-Ren Shyu