arXiv Computer Vision

A comparative assessment of global building and settlement datasets across geographic and settlement contexts

Hugging Face Trending Papers
Jul 22

Global Building Area Estimation Products: How Accurate Are They?

Geo-spatial rasters of building footprint area are useful for a variety of tasks, such as monitoring urbanization, improving energy efficiency, and tracking greenhouse gas emissions. There are now multiple global building raster datasets, however there lacks an independent, comprehensive, and fair assessment of their accuracy.

arXiv Machine Learning
Aug 19

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

The study presents a method for estimating building heights in a large Brazilian city using freely available satellite data, including TerraSAR-X StripMap, PlanetScope, and Sentinel-1. A geographically weighted random forest model achieved an RMSE of 5.34 m and an R² of 0.756 against LiDAR reference data, with local feature importance varying by building type and context. The results highlight that no single sensor dominates across all scenarios, offering guidance for selecting satellite-derived products in different urban settings.

By Guilherme Iablonovski, Pierre-Louis Frison, Tatiana Silva da Silva
Hugging Face Trending Papers
Aug 18

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

The paper presents a method for estimating building heights in a large Brazilian city using freely available satellite data, including TerraSAR‑X StripMap, PlanetScope, and Sentinel‑1. By integrating these sources in a geographically weighted random forest, the authors achieve an RMSE of 5.34 m and an R² of 0.756 against a LiDAR reference. The study also reveals that different predictors dominate in different urban contexts, offering guidance on sensor selection for building‑height mapping.

arXiv Computer Vision
1d ago

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.

By Kaizhen Tan
Hugging Face Trending Papers
Jul 22

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment.

arXiv Computer Vision
5d ago

AlphaEarth distinguishes cities but compresses urban variation

AlphaEarth, a satellite foundation model, maps Earth’s surface into numerical embeddings that allow comparison across places and time. An audit of its representations for 1,000 urban areas in 162 countries shows that cities occupy a distinct but overlapping region on the hypersphere, with continent, climate, and degrees of urbanisation explaining a portion of the variation. The study finds that cities in developing countries exhibit less contrast in vegetation and texture, and that annual changes in a city’s representation are largely driven by model updates rather than pixel changes.

By Andrew Renninger
arXiv Machine Learning
4d ago

DeepC4: Deep Conditional Census-Constrained Clustering for Large-scale Multitask Spatial Disaggregation of Urban Morphology

DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.

By Joshua Dimasaka, Christian Gei{\ss}, Emily So
arXiv Computer Vision
Sep 1

Multi-Sensor Mapping of Vulnerable Urban Settlements Using SAR, Multispectral, and Hyperspectral Imagery: A Case Study in C\'ordoba, Argentina

The paper introduces a multi‑sensor deep learning framework for mapping informal settlements in Córdoba, Argentina, using high‑resolution PlanetScope multispectral imagery, COSMO‑SkyMed SAR data, and medium‑resolution PRISMA hyperspectral observations. It compares SAR‑only, MS‑only, and various fusion strategies (early, middle, late) and finds that late fusion with hyperspectral data (LF+HS) delivers the best balance of classification accuracy and spatial precision. The study also demonstrates that detections outside official polygons align with broader municipal vulnerability layers and that identified settlements show higher surface temperatures during a heatwave, highlighting localized heat amplification.

By Luigi Russo, Anabella Ferral, Silvia Liberata Ullo, Paolo Gamba
arXiv Machine Learning
Aug 21

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.

By Joan Perez, Giovanni Fusco
arXiv Machine Learning
Jun 19

Exploring the potential of AlphaEarth and TESSERA embeddings for Fine-scale Local Climate Zone Mapping: A case study across five cities in Switzerland

arXiv:2606. 20034v1 Announce Type: new Abstract: Understanding urban spatial morphology is critical for climate modeling, risk assessment, and sustainable urban design, and Local Climate Zone (LCZ) mapping provides the basic framework for this.

By Htet Yamin Ko Ko, Clement Atzberger
arXiv Computer Vision
Sep 4

Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications

The paper introduces GBA-GCs, a large-scale multimodal benchmark for identifying gated and open residential compounds in China’s Greater Bay Area, comprising 37,444 compounds with satellite imagery, metadata, and verified labels. It presents MCGC, a vision-centric multimodal framework that fuses imagery, text, and structured data to accurately classify gated communities, outperforming existing baselines. Using the model, the authors map gated communities across the metropolitan area and uncover equity-related patterns such as clustered gated zones, privatized green space, and diminished pedestrian connectivity.

By Minwei Zhao, Weiming Zhang, Jiawang Du, Qiming Liu, Weiming Zhuang, Pei Nie, Cai Wu