arXiv Computer Vision

Seeing the City or Recognizing the Place? What Street-View Imagery Adds Beyond Existing Urban Data in VLM Urban Sensing

The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.

arXiv Machine Learning
Aug 21

From Street View Imagery to Street Quality Indicators: Vision Language Inference for the Suburban 15-minute City

arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.

By Joan Perez, Giovanni Fusco
arXiv Computer Vision
Sep 2

You Cannot Photograph the Same Street Twice: Reliability Limits in Vision-Language Measurement of Urban Change

Vision‑language models used to gauge urban change from repeated street‑level images exhibit limited reliability at single locations. In a study of 4,648 image pairs from 435 Google Street View points across five U.S. cities, re‑photographing the same street altered perception scores by an average of 0.80 points—about two‑thirds of the difference between distinct streets—while repeated model calls added negligible variation. Although image re‑encoding, prompt order, and various image statistics contributed modestly, a small systematic drift (~0.1 points) persisted and grew with time between captures, suggesting minor unrecorded physical changes. Controlled experiments revealed that varying camera and image properties can shift scores, and that camera geometry alone caused a model to falsely report change in 45% of identical scenes; normalising to a common virtual camera reduced this to 7.5%. Despite these individual‑point unreliabilities, aggregating many paired observations recovers a clear redevelopment signal, indicating that such models are dependable at large scales but not for single‑location assessments.

By Kaizhen Tan
Hugging Face Trending Papers
Jul 22

How Does Urban Context Relate to Residential Building Health? A Vision-POI Fusion Framework for Building-Level Housing Inspection

Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment.

arXiv Computer Vision
Aug 27

OpenCVL: An Open, Diverse, and Large-Scale Dataset for Fine-Grained Cross-View Localization

OpenCVL is a large, open dataset for fine-grained cross-view localization, comprising 617,388 ground‑aerial image pairs from 41 European cities. It blends high‑end sensor data with diverse in‑the‑wild images and includes a curation framework to correct pose annotations, enabling reliable evaluation. The dataset also offers cross‑area and snowy test sets to probe generalization, and experiments show that adding noisy in‑the‑wild data improves model performance on clean tests.

By Zimin Xia, Mubariz Zaffar, Junsheng Fu, Alexandre Alahi, Julian F. P. Kooij
arXiv Computer Vision
Sep 24

A comparative assessment of global building and settlement datasets across geographic and settlement contexts

arXiv:2609.28154v1 Announce Type: new Abstract: Global building and settlement datasets increasingly support population mapping, exposure assessment, urban monitoring, and other analyses of the built...

By Rufai Omowunmi Balogun, Caroline Margaux Gevaert, Capucine Riom, Derrick Mirindi, Aaron Opdyke, Hamed Alemohammad, Pierre Chrzanowski, Edward Charles Anderson
Hugging Face Trending Papers
Aug 17

Remote-Sensing City Layout Extraction with MLLM

Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules.

arXiv Computer Vision
Sep 1

SVI2LoD3: Agent-Driven Reconstruction of LoD3 Facade Openings in Semantic 3D City Models from Volunteered Street View Imagery using Large Language and Visual Models

arXiv:2608.29992v1 Announce Type: new Abstract: This paper presents an end-to-end, agent-driven pipeline for the LoD3 reconstruction of facade openings in 3D city models, producing directly usable Ci...

By Elmehdi Kanna, Lukas Arzoumanidis, Huynh Duc An Son Nguyen, Youness Dehbi
arXiv Machine Learning
Aug 19

Spatially explicit feature importance for building height estimation using research-access high-resolution SAR and optical sensors

The study presents a method for estimating building heights in a large Brazilian city using freely available satellite data, including TerraSAR-X StripMap, PlanetScope, and Sentinel-1. A geographically weighted random forest model achieved an RMSE of 5.34 m and an R² of 0.756 against LiDAR reference data, with local feature importance varying by building type and context. The results highlight that no single sensor dominates across all scenarios, offering guidance for selecting satellite-derived products in different urban settings.

By Guilherme Iablonovski, Pierre-Louis Frison, Tatiana Silva da Silva