Housing-level urban physical examination is essential for identifying residential building problems and supporting targeted urban renewal. Existing automated inspection studies primarily rely on individual images and rarely examine whether surrounding urban functional context can provide supplementary information for building-level assessment.
arXiv:2606. 15890v1 Announce Type: new Abstract: Understanding urban wellbeing from multimodal data requires integrating heterogeneous spatial and temporal signals, posing significant challenges for current multimodal large language models (MLLMs).
By Yanxin Xi, Xiang Su, Jie Feng, Yu Liu, Sasu Tarkoma, Pan Hui
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
The paper introduces an adaptive region‑dividing strategy that projects a 3D point cloud onto a bird’s‑eye‑view plane to detect building regions, then back‑projects bounding boxes to create structure‑aligned training blocks for unified scene‑level evaluation. It also proposes a fine‑grained classification model using a point transformer classifier and a spatially‑supervised contrastive loss to improve inter‑class discriminability, addressing class imbalance with a weighted cross‑entropy. Experiments on UrbanBIS and STPLS3D datasets show the method outperforms state‑of‑the‑art approaches in both building instance segmentation and fine‑grained classification.
By Weiyuan Zhang, Qi Zhang, Hui Huang
DeepC4 is a deep learning-based spatial disaggregation method that uses local census statistics as cluster-level constraints and incorporates multiple conditional label relationships in a multitask learning framework. Applied to Rwandan urban morphology, it achieves macro‑F1 scores of 0.63, 0.78, and 0.45 for roof, wall, and height prediction, respectively, and estimates national dwelling and occupant counts within about 1.1% error compared to census records. The approach outperforms existing GEM and METEOR methods and covers 32‑49% more 500‑meter grid pixels across provinces.
By Joshua Dimasaka, Christian Gei{\ss}, Emily So
The paper presents a multimodal machine learning framework that classifies Emirati residential architectural styles by combining visual features from images and textual descriptions using OpenAI's CLIP model. The unified 512‑dimensional embeddings are reduced with UMAP, clustered with K‑Means, and then used to train an SVM classifier, achieving a 98% accuracy across eight style clusters. This approach outperforms previous studies and demonstrates the value of integrating visual and textual data for cultural heritage analysis.
By Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir