MemeCULT-1K: Benchmarking South Asian Cultural Context and Humor Understanding of Multimodal Models
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
BanglaVerse is a new benchmark that evaluates multilingual vision‑language models on Bengali culture, covering nine visual domains and expanding to four languages and five Bangla dialects for a total of about 32,200 artifacts. It includes visual question answering and captioning tasks built from 1,152 manually curated images. Experiments show that models perform worse on dialectal variants and that missing cultural knowledge, rather than visual grounding, is the main bottleneck.
arXiv:2609.00491v1 Announce Type: new Abstract: Communicating across cultures is inherently challenging, especially through culturally dense and ambiguous formats like memes. While people expect larg...
arXiv:2607. 03981v1 Announce Type: cross Abstract: Memes have become influential communication tools on social media, combining viral visuals with concise messaging to convey impactful ideas.
arXiv:2508. 05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings.
The study investigates how large language models (LLMs) handle diverse Indian oral traditions, using the Rajasthani Pabuji epic, Tamil Sangam poetry, and Bengali folk tales as case studies. By prompting Claude Sonnet and Gemini with 54 generation requests across generic, culturally specific, and regional-language prompts, the authors measured reference drift and cross-tradition convergence using Sentence‑BERT embeddings. Results show that while outputs stay closer to their own tradition than to others, there is significant cross‑tradition similarity (0.52–0.66), indicating partial homogenisation; moreover, regional‑language prompting consistently reduced fidelity to authentic traditions.
The paper introduces CMPM, a Chinese Multi-Panel Meme benchmark comprising 1,214 annotated samples that capture five structural types, ordering dependencies, panel-order constraints, and optional comment context. It defines a two-layer evaluation: Task 1 tests structure typing and order-sensitive panel sequencing, while Task 2 assesses Chinese meme explanation generation using human ratings across visual, panel, humor, context, and faithfulness dimensions. Benchmarking five large vision‑language models shows that accuracy on canonical displays does not guarantee order understanding, as performance drops sharply under shuffled conditions, and that Gemini 3.1 Pro and GPT‑5.5 outperform open models in Task 2, with comment context providing only modest gains.