arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.
By Joan Perez, Giovanni Fusco
arXiv:2607. 14756v1 Announce Type: new Abstract: This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images.
By Zahratu Shabrina, Muhammad Asa, Jin Rui, Lu Yin, Stephen Law
Vision‑language models used to gauge urban change from repeated street‑level images exhibit limited reliability at single locations. In a study of 4,648 image pairs from 435 Google Street View points across five U.S. cities, re‑photographing the same street altered perception scores by an average of 0.80 points—about two‑thirds of the difference between distinct streets—while repeated model calls added negligible variation. Although image re‑encoding, prompt order, and various image statistics contributed modestly, a small systematic drift (~0.1 points) persisted and grew with time between captures, suggesting minor unrecorded physical changes. Controlled experiments revealed that varying camera and image properties can shift scores, and that camera geometry alone caused a model to falsely report change in 45% of identical scenes; normalising to a common virtual camera reduced this to 7.5%. Despite these individual‑point unreliabilities, aggregating many paired observations recovers a clear redevelopment signal, indicating that such models are dependable at large scales but not for single‑location assessments.
By Kaizhen Tan
arXiv:2503.00610v2 Announce Type: replace-cross
Abstract: Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult t...
By Ciro Beneduce, Bruno Lepri, Massimiliano Luca
arXiv:2610.01870v1 Announce Type: new
Abstract: Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventi...
By Hosam Elgendy, Utkarsh Mall
arXiv:2608. 06934v1 Announce Type: new Abstract: Visual perception of walkability varies substantially across individuals, reflecting differences in personal characteristics, experiences, and preferences.
By Moloud Damandeh, Meead Saberi
arXiv:2608. 14663v1 Announce Type: new Abstract: As global populations age, enhancing neighborhood walkability through inclusive urban design is important for mitigating built environment (BE) barriers that discourage physical activity and social participation among older adults.
By Houhao Liang, Kresimir Friganovic, Joanne Kua, Noor Hafizah Ismail, Su Su, Bryan Yijia Tan, Navrag B. Singh, Panos Mavros
PlaceSeek is a human‑centered geospatial retrieval framework that maps natural‑language queries to street‑view images by decomposing queries into functional and affective sub‑intents. It uses a Semantic Grounding Module to verify that candidate images contain the physical evidence needed for the intended activity, and an Affective Alignment Module to re‑rank these candidates based on human urban perception judgments. Evaluated on 31,956 Milan street‑view locations, PlaceSeek achieves high precision and ranking metrics, outperforming several vision‑language baselines and demonstrating the importance of both physical grounding and affective alignment for complex urban spatial queries.
By Ziqi Cui, Shangyu Lou
The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.
By Kaizhen Tan
Urban environments are shaped by design choices with long-term implications for health, safety, and quality of life, yet evaluating proposed interventions remains costly, time-consuming, and often imp...
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
The paper introduces a vision‑language framework that extracts social indicators from street‑level imagery, converting panoramic views into sidewalk‑facing sideviews with timestamps. Using a VLM‑based activity detection system, it codes each pedestrian across ten observable dimensions, producing a Social Dwelling Index (SDI) that captures grouping, dwelling, activity diversity, and accessibility flags. Applied to over 100,000 sideviews in New York City, the study finds that pedestrian volume and SDI are only weakly correlated, indicating that high foot traffic does not necessarily equate to intense social activity.
By Liu Liu, Andres Sevtsuk