arXiv Computation and Language By Nurul Labib Sayeedi, Md. Faiyaz Abdullah Sayeedi, Shubhashis Roy Dipta, Mahbub E Sobhani, Rubaya Tabassum, Ariful Ekraj Hridoy, Mehraj Mahmood, Md. Tarek Hasan, Swakkhar Shatabda

Many Dialects, Many Languages, One Cultural Lens: Evaluating Multilingual VLMs for Bengali Culture Understanding Across Historically Linked Languages and Regional Dialects

Read the original on arXiv Computation and Language →

BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
1d ago

MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models

arXiv:2609.01772v1 Announce Type: new Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vi...

By Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed
Hugging Face Trending Papers
Jun 8

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.

arXiv AI
3d ago

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, cove...

By Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam
arXiv Computation and Language
Aug 25

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understa...

By S{\l}awomir Dadas, Micha{\l} Pere{\l}kiewicz, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Bart{\l}omiej Jaworski, Izabela Wo\'zniakowska