arXiv AI By Ahmed Ammar Kubba, Manar Abu Talib, Iman Ibrahim, Qassim Nasir

Multimodal Cultural Heritage Architectural Style Classification for Residential Buildings in the UAE Based on CLIP Embeddings and SVM

Read the original on arXiv AI →

The paper presents a multimodal machine learning framework that classifies Emirati residential architectural styles by combining visual features from images and textual descriptions using OpenAI's CLIP model. The unified 512‑dimensional embeddings are reduced with UMAP, clustered with K‑Means, and then used to train an SVM classifier, achieving a 98% accuracy across eight style clusters. This approach outperforms previous studies and demonstrates the value of integrating visual and textual data for cultural heritage analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 8

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.

arXiv Computer Vision
Sep 4

Urban Boundaries, Social Barriers: A Benchmark and Vision-Centric Framework for Mapping Gated Communities and Equity Implications

The paper introduces GBA-GCs, a large-scale multimodal benchmark for identifying gated and open residential compounds in China’s Greater Bay Area, comprising 37,444 compounds with satellite imagery, metadata, and verified labels. It presents MCGC, a vision-centric multimodal framework that fuses imagery, text, and structured data to accurately classify gated communities, outperforming existing baselines. Using the model, the authors map gated communities across the metropolitan area and uncover equity-related patterns such as clustered gated zones, privatized green space, and diminished pedestrian connectivity.

By Minwei Zhao, Weiming Zhang, Jiawang Du, Qiming Liu, Weiming Zhuang, Pei Nie, Cai Wu
arXiv AI
Sep 16

CLIP Embeddings for AI-Generated Image Detection: A Few-Shot Study with Lightweight Classifier

The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.

By Ziyang Ou
arXiv Machine Learning
Aug 28

How AI Experiences Art: Emergent Aesthetic Structure in a Self-Supervised Multimodal Embedding Space

The paper introduces a self‑supervised framework that maps text, audio, image, and video into a shared 256‑dimensional embedding space and uses iterative clustering to uncover aesthetic structure. It examines how AI’s cluster assignments diverge from human affective labels on a weakly supervised multimodal dataset. The study highlights implications for cross‑modal similarity, media organization for Retrieval‑Augmented Generation, and automated data labeling.

By Corey D. C. Heath