Hugging Face Trending Papers

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.

arXiv Computation and Language
Aug 28

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.

By Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Mahbub E Sobhani, Rubaya Tabassum, Ariful Ekraj Hridoy, Mehraj Mahmood, Md. Tarek Hasan, Swakkhar Shatabda
arXiv AI
Sep 17

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.

By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
arXiv Machine Learning
Jun 5

Almieyar-Oryx-BloomBench: A Bilingual Multimodal Benchmark for Cognitively Informed Evaluation of Vision-Language Models

arXiv:2606. 05531v1 Announce Type: cross Abstract: Despite the rapid progress of Vision-Language Models (VLMs), the field lacks benchmarks that rigorously diagnose their true reasoning abilities and chart meaningful progress toward human-like multimodal intelligence.

By Mohammad Mahdi Abootorabi, Omid Ghahroodi, Anas Madkoor, Marzia Nouri, Doratossadat Dastgheib, Mohamed Hefeeda, Ehsaneddin Asgari
arXiv AI
Sep 1

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, cove...

By Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam
arXiv AI
Sep 10

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

NormViz introduces a new benchmark, NormViz‑Bench, comprising 3,268 contrastive image pairs from 16 countries that test AI’s ability to recognize culturally relevant visual norms. Each pair differs only in a behavior that changes its cultural interpretation, and images are labeled as conforming, violating, or irrelevant to local norms, requiring both images to be correctly classified. The benchmark shows current VLMs perform poorly, and a complementary training set, NormViz‑Train, offers a path to improve performance by teaching models to link visual perception with cultural significance.

By Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja, Abhinav Sukumar Rao, Ashima Suvarna, Nanyun Peng, Saadia Gabriel, Maarten Sap