The paper presents a multimodal machine learning framework that classifies Emirati residential architectural styles by combining visual features from images and textual descriptions using OpenAI's CLIP model. The unified 512‑dimensional embeddings are reduced with UMAP, clustered with K‑Means, and then used to train an SVM classifier, achieving a 98% accuracy across eight style clusters. This approach outperforms previous studies and demonstrates the value of integrating visual and textual data for cultural heritage analysis.
By Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir
arXiv:2509. 09794v5 Announce Type: replace Abstract: Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems.
By Jackson Eshbaugh, Chetan Tiwari, Jorge Silveyra
arXiv:2503.00610v2 Announce Type: replace-cross
Abstract: Understanding how people perceive urban environments is essential for inclusive planning, yet conventional surveys are costly and difficult t...
By Ciro Beneduce, Bruno Lepri, Massimiliano Luca
Remote-sensing systems usually describe urban content with detection boxes, semantic masks, or vector boundaries. Such outputs locate classes and support image-plane scoring, yet they do not by themselves constitute an executable layout that retains object identities, typed relations, topology, and regeneration rules.
The study evaluates how much street‑view imagery contributes to urban attribute prediction beyond existing public data. By comparing image‑based models with seven attributes from five public sources and three vision‑language models, the authors find that images outperform other data for building type, function, and low‑rise floor count, while existing data match or exceed image performance for road damage, curb ramps, and house price. The benefit of images varies with visual legibility and local data coverage, suggesting that image value depends on how well the scene is captured and how much complementary data is available.
By Kaizhen Tan
arXiv:2606. 00871v1 Announce Type: cross Abstract: Vision-language models (VLMs) are increasingly used to generate structured descriptions of street-level imagery for tasks such as streetscape auditing, mapping, and public consultation.
By Rashid Mushkani