AVMeme Exam: A Multimodal Multilingual Multicultural Benchmark for LLMs' Contextual and Cultural Knowledge and Thinking
Read the original on arXiv Computation and Language →AVMeme Exam is a human‑curated benchmark featuring over a thousand iconic Internet audio‑visual clips—including speech, songs, music, and sound effects—each paired with a unique Q&A that probes understanding from surface content to context, emotion, usage, and world knowledge. The benchmark also provides metadata such as original year, transcript, summary, and sensitivity. Evaluations of state‑of‑the‑art multimodal large language models (MLLMs) and human participants reveal that current models perform poorly on textless music and sound effects and struggle to think in cultural and contextual terms compared to surface content.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.