arXiv AI

CulturalMenuBench: Probing the Knowledge-Application Gap in Multimodal Culinary Reasoning

CulturalMenuBench is a new benchmark comprising 4,870 culinary items in 10 languages across 18 regions, designed to test multimodal language models on tasks that combine dish recognition, step-by-step cooking images, ingredients, procedural text, and regional labels. The benchmark reveals a large knowledge‑application gap: models that score over 94% on standard multiple‑choice questions fall to at most 56% when attributing dishes to Chinese regional cuisines, indicating that cultural knowledge is present but not activated by visual input. Diagnostic analyses show that accuracy is driven by visual distinctiveness rather than cultural structure, and that removing sequential cooking images selectively harms process‑grounded tasks, confirming the need for procedural evidence.

arXiv AI
Aug 5

OliveGemma: A 3 Billion Visual Language Model for Recognising the Mediterranean & European Diet

arXiv:2608. 03428v1 Announce Type: cross Abstract: Image based dietary assessment offers a scalable alternative to self reported food diaries, yet fine-grained food recognition remains challenging due to high intra-class variability and visually similar dishes.

By Dimitrios I. Zaridis, Traianos Tsiokris, Vasileios C. Pezoulas, Daphni Plati, Eugenia Mylona, Eleni Georga, Nikos Tsiknakis, Antonis Sakellarios, Dimitrios I. Fotiadis
arXiv AI
Sep 10

NormViz: A Benchmark and Framework for Grounding Multimodal Reasoning in Global Cultures

NormViz introduces a new benchmark, NormViz‑Bench, comprising 3,268 contrastive image pairs from 16 countries that test AI’s ability to recognize culturally relevant visual norms. Each pair differs only in a behavior that changes its cultural interpretation, and images are labeled as conforming, violating, or irrelevant to local norms, requiring both images to be correctly classified. The benchmark shows current VLMs perform poorly, and a complementary training set, NormViz‑Train, offers a path to improve performance by teaching models to link visual perception with cultural significance.

By Akhila Yerukola, Fabrice Y Harel-Canada, Simran Khanuja, Abhinav Sukumar Rao, Ashima Suvarna, Nanyun Peng, Saadia Gabriel, Maarten Sap
arXiv Computation and Language
Aug 28

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

CultureVidBench is a new benchmark that evaluates how well text‑to‑video generation models capture cultural details. It contains 1,000 prompts spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects, grouped into material culture, social practice & performance, and ritual & ceremony. Human studies and automated assessments show that while current models perform well on semantic adherence and visual quality, they often miss fine‑grained cultural details, especially for underrepresented regions and multimodal cues.

By Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
Hugging Face Trending Papers
Jun 8

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.

arXiv AI
Sep 17

MUSE: Benchmarking Large Vision-Language Models on Multi-Modal Understanding in Situated Education

MUSE is a new benchmark designed to evaluate large vision‑language models on artistic image understanding within situated educational contexts. It separates image annotation from question generation, offering twelve tasks that cover visual perception, semantic and affective interpretation, cultural understanding, and compositional reasoning across diverse artistic images from Singaporean, Southeast Asian, and Western traditions. The benchmark reveals significant gaps in model performance, especially in affective interpretation and compositional reasoning, and highlights common failure modes for trustworthy educational multimodal systems.

By Luyao Zhu, Xun Wei Yee, Wei Li, Mun Thye Mak, Wee Siong Ng
arXiv AI
Aug 25

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

The Cultural Moment Benchmark (CMB) evaluates video cultural reasoning in Southeast Asia by testing three distinct abilities: naming a cultural concept, visually recognizing it in a video, and locating its sub‑events in time. It contains 306 expert‑curated concepts from seven countries across five categories, with each concept assessed through three stages that use semantic‑similarity distractors, unlabeled video moments, and free‑form temporal localization. Experiments on six vision‑language models reveal varied failure modes, limited cascading between abilities, and differing impacts of audio and subtitles, while a human study shows even experts struggle with concepts from neighboring countries.

By Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
arXiv Computation and Language
Sep 3

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

MemeCULT-1K is a multilingual benchmark of 1,000 South Asian memes in Bengali, English, and Hindi, each paired with a cultural context note and three human-written explanations, plus an additional set of 54 Bengali regional dialect memes. The study evaluates thirteen vision‑language models under meme‑only and context‑aware settings, showing that providing minimal cultural context consistently improves performance across all models and languages. Error analysis indicates closed‑source models struggle with entity and reference misidentification, while open‑source models are limited by broader cultural knowledge gaps, especially in linguistic and phonological aspects.

By Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed
arXiv AI
Jul 7

HCSU: A Dataset and Benchmark for Fine-Grained Historical Calligraphy Style Understanding

arXiv:2607. 04147v1 Announce Type: cross Abstract: Automated fine-grained perception of calligraphy styles--a task vital to cultural heritage preservation--remains a critical challenge for Large Vision-Language Models (LVLMs), largely constrained by existing datasets that suffer from modal mixture and flattened labels.

By Yinsheng Yao, Yan Liu, Chen Ye