arXiv:2608. 20026v1 Announce Type: cross Abstract: Streetscape quality has become a central concern in contemporary urban planning, particularly within the framework of the pedestrian-friendly 15-minute city, where walkability and public-space quality are increasingly recognized as key determinants of urban performance.
By Joan Perez, Giovanni Fusco
arXiv:2606. 00404v1 Announce Type: cross Abstract: Implicit neural representations (INRs) model a signal as a continuous coordinate-to-value function.
By Haoan Feng, Xin Xu, Leila De Floriani
The paper presents a method for enhancing coarse 5 m digital surface models (DSMs) to 0.5 m resolution by guiding a denoising diffusion process with high‑resolution spectral images. This approach transfers fine visual details—such as crisp outlines and roof structures—from the imagery into the elevation maps, yielding more accurate surface geometry than traditional interpolation or filtering. Experiments on Central European cities show that the resulting DSMs exhibit improved structural detail and overall quality.
arXiv:2607. 22342v1 Announce Type: new Abstract: The development of effective urban climate adaptation strategies requires comprehensive spatial information on rooftops and buildings, since such information underpins the assessment of ecosystem services provided by green infrastructure, particularly for urban heat island (UHI) mitigation.
By Htet Yamin Ko Ko
arXiv:2607. 14756v1 Announce Type: new Abstract: This research investigates the potential of Vision-Language Models (VLMs) to infer building typologies: Construction, Current Use, and Storeys from Google Street View (GSV) images.
By Zahratu Shabrina, Muhammad Asa, Jin Rui, Lu Yin, Stephen Law
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
The paper proposes a method to enhance coarse 5 m digital surface models (DSMs) to 0.5 m resolution by guiding the super‑resolution process with high‑resolution spectral images. It uses denoising diffusion to transfer image‑visible details, such as crisp outlines and roof structures, into the elevation maps, achieving more accurate surface geometry than traditional interpolation or filtering. Experiments on Central European cities show that the approach yields high‑quality DSMs with improved structural detail.
By Armand Mihai Nicolicioiu, Dominik Narnhofer, Nando Metzger, Daniel Panangian, Ksenia Bittner, Konrad Schindler
The paper proposes a method to adapt the Depth Anything V2 (DAV2) zero‑shot relative depth model for estimating lunar surface height. By fine‑tuning DAV2 with publicly available stereophotogrammetry‑derived DEM data, the authors achieve a significant performance boost over the unadapted zero‑shot model. This improved estimator can provide more accurate relative height information useful for hazard detection in future ESA lunar landings.
By Patrick Bauer, Marius Schwinning, Melanie Siegel, Andreas Weinmann, Hichem Snoussi
BEACON is a tri‑modal contrastive learning framework that enriches AlphaEarth embeddings by aligning physical representations from Earth‑observation imagery with semantic POI text and human behavioral POI visitation data, while keeping the deployed model image‑only. In a Houston case study, BEACON outperformed six baselines on nine downstream tasks, achieving up to 43% higher R² for obesity prevalence, 34% for poor mental health, and 22% for median household income under a linear probe.
By Hao Tian, Heng Cai, Yifan Yang
ImplicitTerrainV2 introduces a wavelet-guided, spatially adaptive neural representation for digital elevation models (DEMs). It uses a wavelet complexity field to localize high-frequency capacity to complex terrain, adaptive sampling to focus training, and gradient matching to preserve smooth manifold structure. After mixed-precision quantization and entropy coding, the model achieves 1.23 bpp with only a 0.28 dB PSNR loss, outperforming prior work by 5.70 dB while using 3.2× fewer parameters and training in 55 s per tile on a single GPU.
By Haoan Feng, Xin Xu, Leila De Floriani
arXiv:2608.29426v1 Announce Type: cross
Abstract: Reliable semantic representations derived from city-scale 3D models are increasingly important for urban analysis, infrastructure monitoring, autonom...
By Alexander Rusnak, Sophia Kovalenko, Jingru Wang, Ismail Moudden, Xiru Wang, Fr\'ed\'eric Kaplan
Vision‑language models used to gauge urban change from repeated street‑level images exhibit limited reliability at single locations. In a study of 4,648 image pairs from 435 Google Street View points across five U.S. cities, re‑photographing the same street altered perception scores by an average of 0.80 points—about two‑thirds of the difference between distinct streets—while repeated model calls added negligible variation. Although image re‑encoding, prompt order, and various image statistics contributed modestly, a small systematic drift (~0.1 points) persisted and grew with time between captures, suggesting minor unrecorded physical changes. Controlled experiments revealed that varying camera and image properties can shift scores, and that camera geometry alone caused a model to falsely report change in 45% of identical scenes; normalising to a common virtual camera reduced this to 7.5%. Despite these individual‑point unreliabilities, aggregating many paired observations recovers a clear redevelopment signal, indicating that such models are dependable at large scales but not for single‑location assessments.
By Kaizhen Tan