arXiv AI

Platonic Representations for Poverty Mapping: Unified Vision-Language Codes or Agent-Induced Novelty?

arXiv:2508. 01109v3 Announce Type: replace Abstract: We investigate whether socioeconomic indicators, like household wealth, leave recoverable informational imprints in both satellite imagery (capturing features like buildings and roads) and Internet-sourced text (reflecting historical, cultural, and narratives of neighborhoods).

arXiv AI
Aug 19

MoRA: Mobility as the Backbone for Geospatial Representation Learning at Scale

MoRA is a human‑centric geospatial representation learning framework that uses a large mobility graph as its backbone to fuse spatial tokenization, graph neural networks, and asymmetric contrastive learning. It aligns over 100 million points of interest, massive remote sensing imagery, and structured demographic data with a billion‑edge mobility graph, producing compact 128‑dimensional embeddings that capture socio‑economic context and functional roles of locations. On a benchmark of nine downstream social and economic prediction tasks, MoRA outperforms state‑of‑the‑art models by an average of 12.9% and demonstrates scaling behavior analogous to large language models.

By Ya Wen, Jixuan Cai, Qiyao Ma, Linyan Li, Xinhua Chen, Chris Webster, Yulun Zhou
arXiv Machine Learning
5d ago

DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology

DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.

By Joshua Dimasaka, Christian Gei{\ss}, Emily So
arXiv AI
Jun 8

Textual Supervision Enhances Geospatial Representations in Vision-Language Models

arXiv:2606. 07172v1 Announce Type: cross Abstract: Geospatial understanding is a critical yet underexplored dimension in the development of machine learning systems for tasks such as image geolocation and spatial reasoning.

By Marcelo Sartori Locatelli, Fernando Tonucci, Jea Kwon, Luiz Felipe Vecchietti, Bryan Nathanael Wijaya, Cheng Yaw Low, Virgilio Almeida, Meeyoung Cha
Hugging Face Trending Papers
Aug 4

Geo-Embed: Towards Unified Multimodal Embeddings for Urban Understanding

Geospatial and urban applications increasingly require models to compare heterogeneous evidence across street-view imagery, remote-sensing observations, text descriptions, region proposals, and temporal change cues. However, existing multimodal embedding models and benchmarks are still largely designed and evaluated around general-purpose image-text matching, leaving unclear whether unified embedding space can support heterogeneous geospatial tasks involving spatial relationships, fine-grained semantics, and temporal changes.