The paper presents a multimodal machine learning framework that classifies Emirati residential architectural styles by combining visual features from images and textual descriptions using OpenAI's CLIP model. The unified 512‑dimensional embeddings are reduced with UMAP, clustered with K‑Means, and then used to train an SVM classifier, achieving a 98% accuracy across eight style clusters. This approach outperforms previous studies and demonstrates the value of integrating visual and textual data for cultural heritage analysis.
By Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir
The paper introduces an adaptive region‑dividing strategy that projects a 3D point cloud onto a bird’s‑eye‑view plane to detect building regions, then back‑projects bounding boxes to create structure‑aligned training blocks for unified scene‑level evaluation. It also proposes a fine‑grained classification model using a point transformer classifier and a spatially‑supervised contrastive loss to improve inter‑class discriminability, addressing class imbalance with a weighted cross‑entropy. Experiments on UrbanBIS and STPLS3D datasets show the method outperforms state‑of‑the‑art approaches in both building instance segmentation and fine‑grained classification.
By Weiyuan Zhang, Qi Zhang, Hui Huang
arXiv:2608.03323v2 Announce Type: replace
Abstract: Estimating room layouts from multi-view imagery is a core task for indoor scene understanding. Existing methods are typically limited either by poo...
By Gustav Hanning, Shaohui Liu, R\'emi Pautrat, Marc Pollefeys, Kalle {\AA}str\"om, Viktor Larsson
arXiv:2512. 04832v3 Announce Type: replace-cross Abstract: We present a BIM-native tokenization for room-level layout synthesis in Building Information Modeling (BIM) scenes.
By Manuel Ladron de Guevara, Jinmo Rhee, Ardavan Bidgoli, Vaidas Razgaitis, Michael Bergin
arXiv:2607. 00090v1 Announce Type: cross Abstract: Urban-scale Visual Place Recognition (VPR) aims to identify the geographic location of a query image by matching it against a geo-tagged database.
By Zhiyao Shu, Jiacheng Yang, Yang Lu, Waishan Qiu, Chuan Li, Da Chen
arXiv:2505.12254v3 Announce Type: replace-cross
Abstract: Existing visual place recognition (VPR) datasets predominantly rely on vehicle-mounted imagery, offer limited multimodal diversity, and under...
By Yiwei Ou, Xiaobin Ren, Ronggui Sun, Guansong Gao, Kaiqi Zhao, Manfredo Manfredini
arXiv:2608.31107v1 Announce Type: new
Abstract: The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalizatio...
By Lucas Wojcik, Gabriel E. Lima, Sergio M. Silva Jr., Eduil Nascimento Jr., David Menotti
arXiv:2609.36648v1 Announce Type: new
Abstract: Vision-language pre-training has reshaped image clustering, giving rise to language-assisted image clustering (LaIC), which leverages textual semantics...
By Yuanwei Hu, Bo Peng, Yuheng Jia, Xinting Hu, Yadan Luo, Wenjie Zhu
The advent of foundation models have enabled a new era in zero-shot classification. Yet, key challenges persist. Despite their impressive generalization power that leverages the immense pre-training k...
arXiv:2607. 22919v1 Announce Type: cross Abstract: Multimodal embedding spaces in models like CLIP enable powerful capabilities such as semantic similarity retrieval and cross-modal zero-shot classification.
By Joseph Fioresi, Fabian Caba Heilbron, Pankaj Nathani, Mubarak Shah, Kushal Kafle
Unifying image clustering across different clustering scenarios remains challenging due to fundamental gaps among tasks. We introduce a Guideline-Driven Image Clustering Agent, the first universal framework that bridges these gaps through textual guidelines.
arXiv:2509. 09794v5 Announce Type: replace Abstract: Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems.
By Jackson Eshbaugh, Chetan Tiwari, Jorge Silveyra