arXiv Computation and Language

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

CultureVidBench is a new benchmark that evaluates how well text‑to‑video generation models capture cultural details. It contains 1,000 prompts spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects, grouped into material culture, social practice & performance, and ritual & ceremony. Human studies and automated assessments show that while current models perform well on semantic adherence and visual quality, they often miss fine‑grained cultural details, especially for underrepresented regions and multimodal cues.

arXiv AI
Aug 14

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

arXiv:2608. 13210v1 Announce Type: cross Abstract: Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit.

By Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
arXiv AI
Aug 25

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

The Cultural Moment Benchmark (CMB) evaluates video cultural reasoning in Southeast Asia by testing three distinct abilities: naming a cultural concept, visually recognizing it in a video, and locating its sub‑events in time. It contains 306 expert‑curated concepts from seven countries across five categories, with each concept assessed through three stages that use semantic‑similarity distractors, unlabeled video moments, and free‑form temporal localization. Experiments on six vision‑language models reveal varied failure modes, limited cascading between abilities, and differing impacts of audio and subtitles, while a human study shows even experts struggle with concepts from neighboring countries.

By Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
arXiv AI
Jul 20

LVSum: A Benchmark for Timestamp-Aware Long Video Summarization

arXiv:2604. 10024v2 Announce Type: replace-cross Abstract: Long video summarization presents significant challenges for multimodal large language models (MLLMs), particularly in maintaining temporal fidelity over extended durations and producing summaries that are both semantically and temporally grounded.

By Alkesh Patel, Melis Ozyildirim, Ying-Chang Cheng, Ganesh Nagarajan
arXiv Computation and Language
Aug 25

PUMA: A Polish Benchmark for Culturally Grounded Multimodal Understanding

arXiv:2608.21853v1 Announce Type: new Abstract: Large language models are increasingly moving beyond text processing, adding support for other modalities such as images and audio. While text understa...

By S{\l}awomir Dadas, Micha{\l} Pere{\l}kiewicz, Rafa{\l} Po\'swiata, Ma{\l}gorzata Gr\k{e}bowiec, Bart{\l}omiej Jaworski, Izabela Wo\'zniakowska
arXiv Computation and Language
4d ago

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
3d ago

ImageEval 2026: Culturally Grounded Arabic Multimodal Evaluation

arXiv:2608.30475v1 Announce Type: cross Abstract: We present an overview of the ImageEval 2026 shared task on culturally grounded Arabic multimodal evaluation. It includes two tasks: (i) AynVQA, cove...

By Samir Abdaljalil, Hunzalah Hassan Bhatti, Ahlam Bashiti, Farina Amir, Md Arid Hasan, Basel Mousi, Nadir Durrani, Fahim Dalvi, Zien Sheikh Ali, Erchin Serpedin, Hasan Kurban, Mustafa Jarrar, Shammur Absar Chowdhury, Firoj Alam