arXiv AI

OSMGraphCLIP: Learning Global Location Representations from OpenStreetMap Graphs

arXiv:2606. 08046v1 Announce Type: new Abstract: We present OSMGraphCLIP, a CLIP-style geospatial representation model that learns global location embeddings from freely available OpenStreetMap (OSM) data.

arXiv AI
Aug 19

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.

By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv Machine Learning
Aug 19

MoRAX: Mobility-based Representation Augmentation for Geospatial Foundation Models

MoRAX is a lightweight framework that augments geospatial foundation model embeddings with functional structure derived from human mobility data. By incorporating mobility flows, MoRAX preserves the coverage and consistency of existing geospatial models while adding information about functional connectivity among urban regions, enabling zero‑shot deployment in unseen cities. Experiments across four cities in two countries show that the MoRAX teacher model outperforms baseline geospatial models on eight socioeconomic and environmental prediction tasks, and the student model—without direct mobility input—approaches the teacher’s performance.

By Ya Wen, Jixuan Cai, Yulun Zhou, Alec Kirkley
arXiv Machine Learning
Sep 1

BEACON: Behavioral and Semantic Enrichment of AlphaEarth Embeddings through Tri-Modal Contrastive Learning

BEACON is a tri‑modal contrastive learning framework that enriches AlphaEarth embeddings by aligning physical representations from Earth‑observation imagery with semantic POI text and human behavioral POI visitation data, while keeping the deployed model image‑only. In a Houston case study, BEACON outperformed six baselines on nine downstream tasks, achieving up to 43% higher R² for obesity prevalence, 34% for poor mental health, and 22% for median household income under a linear probe.

By Hao Tian, Heng Cai, Yifan Yang
arXiv AI
Jul 9

CarbonCLIP: Enhance Carbon Prediction from Satellite Imagery via Integrated Street-View Semantics and Temporal Context Training

arXiv:2607. 07292v1 Announce Type: cross Abstract: Accurately estimating urban carbon emissions is critical for sustainable urban planning, yet many existing approaches remain difficult to apply consistently across cities due to data-source heterogeneity and the lack of fine-grained semantic-temporal context in remote sensing data.

By Zeru Yang, Fang-Ying Gong, Steve H. L. Yim, Chau Yuen
arXiv Machine Learning
Aug 19

A multi-view contrastive learning framework for spatial embeddings in risk modelling

The paper introduces a multi‑view contrastive learning framework that creates low‑dimensional spatial embeddings by combining satellite imagery and OpenStreetMap data across Europe. These embeddings align with coordinate‑based encodings, allowing any dataset with latitude‑longitude pairs to be enriched with meaningful spatial features without needing the original spatial inputs. In case studies on French real‑estate prices and Belgian flood claim counts, models using the embeddings outperform those using raw coordinates, improving predictive accuracy and territorial risk classification while offering explainable spatial effects.

By Freek Holvoet, Christopher Blier-Wong, Katrien Antonio
arXiv Computer Vision
Aug 27

GTPred: Benchmarking MLLMs for Interpretable Geo-localization and Time-of-capture Prediction

GTPred is a new benchmark for geo‑temporal prediction that evaluates multi‑modal large language models (MLLMs) on 370 images taken across 120 years worldwide. It assesses predictions by matching both the year and a hierarchical location sequence, and includes annotated reasoning chains to test intermediate reasoning. Experiments on 15 MLLMs show that while visual perception is strong, models still lack world knowledge and geo‑temporal reasoning, and that adding temporal data improves location inference.

By Jinnao Li, Tingzhu Chen, Changbo Wang
arXiv Machine Learning
5d ago

DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology

DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.

By Joshua Dimasaka, Christian Gei{\ss}, Emily So