arXiv:2507. 00263v2 Announce Type: replace-cross Abstract: The rapid growth of vacation rental (VR) platforms has led to an increasing volume of property images, often uploaded without structured categorization.
By Vignesh Ram Nithin Kappagantula, Shayan Hassantabar
We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.
arXiv:2509. 09794v5 Announce Type: replace Abstract: Computational models have emerged as powerful tools for multi-scale energy modeling research at the building and urban scale, supporting data-driven analysis across building and urban energy systems.
By Jackson Eshbaugh, Chetan Tiwari, Jorge Silveyra
The paper introduces GBA-GCs, a large-scale multimodal benchmark for identifying gated and open residential compounds in China’s Greater Bay Area, comprising 37,444 compounds with satellite imagery, metadata, and verified labels. It presents MCGC, a vision-centric multimodal framework that fuses imagery, text, and structured data to accurately classify gated communities, outperforming existing baselines. Using the model, the authors map gated communities across the metropolitan area and uncover equity-related patterns such as clustered gated zones, privatized green space, and diminished pedestrian connectivity.
By Minwei Zhao, Weiming Zhang, Jiawang Du, Qiming Liu, Weiming Zhuang, Pei Nie, Cai Wu
The paper explores whether CLIP embeddings can detect AI-generated images by using a frozen CLIP model to extract visual embeddings and training lightweight classifiers on top. On the CIFAKE benchmark, the approach achieves 95% accuracy without language reasoning, and 85% accuracy after few-shot adaptation with 20% of the data. Certain image types, such as wide-angle photographs and oil paintings, remain challenging, highlighting unexplored difficulties in AI-generated image classification.
By Ziyang Ou
The paper introduces a self‑supervised framework that maps text, audio, image, and video into a shared 256‑dimensional embedding space and uses iterative clustering to uncover aesthetic structure. It examines how AI’s cluster assignments diverge from human affective labels on a weakly supervised multimodal dataset. The study highlights implications for cross‑modal similarity, media organization for Retrieval‑Augmented Generation, and automated data labeling.
By Corey D. C. Heath