arXiv AI

VideoNorms: Benchmarking Cultural Awareness of Video Language Models

arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.

arXiv Computation and Language
Aug 28

CultureVidBench: Benchmarking Cultural Understanding in Text-to-Video Generation

CultureVidBench is a new benchmark that evaluates how well text‑to‑video generation models capture cultural details. It contains 1,000 prompts spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects, grouped into material culture, social practice & performance, and ritual & ceremony. Human studies and automated assessments show that while current models perform well on semantic adherence and visual quality, they often miss fine‑grained cultural details, especially for underrepresented regions and multimodal cues.

By Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
arXiv AI
Aug 25

Cultural Moment Benchmark: Evaluating Video Cultural Reasoning and Grounding in Southeast Asia

The Cultural Moment Benchmark (CMB) evaluates video cultural reasoning in Southeast Asia by testing three distinct abilities: naming a cultural concept, visually recognizing it in a video, and locating its sub‑events in time. It contains 306 expert‑curated concepts from seven countries across five categories, with each concept assessed through three stages that use semantic‑similarity distractors, unlabeled video moments, and free‑form temporal localization. Experiments on six vision‑language models reveal varied failure modes, limited cascading between abilities, and differing impacts of audio and subtitles, while a human study shows even experts struggle with concepts from neighboring countries.

By Burak Satar, Zhixin Ma, Cheng Yu-Tong, Huy Hoang Tran, Phuong Anh Nguyen, Chong-Wah Ngo
arXiv Computation and Language
Aug 31

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
Aug 14

NARU: A Benchmark for NARrative Evolution and Cultural Nuance Understanding in Japanese Extreme Long Video

arXiv:2608. 13210v1 Announce Type: cross Abstract: Long-form video understanding encompasses tasks that go beyond retrieving isolated events, including tracking an evolving narrative and interpreting social meaning that may remain implicit.

By Yuheng Huang, Jianlang Chen, Jiayang Song, Hua Qi, Aza Kai, Vincent Markert, Edison Marrese-Taylor, Jianjun Zhao, Lei Ma
arXiv Computer Vision
Sep 24

Can Vision-Language Models Analyze Human-Centered Video? Mapping Model Capabilities and Human-AI Collaborative Workflows

The paper investigates when vision‑language models (VLMs) can independently analyze human‑centered video and when human oversight is still needed. By reviewing 1,702 CHI 2026 papers, the authors develop a five‑dimensional taxonomy of video annotation tasks and build a benchmark of 15 representative tasks. Experiments show that VLMs alone achieve near‑human accuracy (HNS = 97.0), while human verification of VLM outputs yields the highest accuracy (HNS = 121.5) and significantly reduces annotation time and cost.

By Xiyuan Shen, Jiuyang Lyu, Seokhyun Hwang, Huanfen Yao, Shwetak Patel, Zhihan Zhang, Jacob O. Wobbrock
Hugging Face Trending Papers
Jun 8

ChinaHeritaQA: A Culturally-Grounded Visual Question Answering Dataset for World Heritage Sites in China

We introduce ChinaHeritaQA, a multimodal benchmark dataset for evaluating the cultural reasoning abilities of vision-language models (VLMs) on UNESCO World Heritage sites in China. The dataset comprises 2,279 in-the-wild images paired with 14,133 bilingual (Chinese/English) multiple-choice QA pairs spanning seven cognitive dimensions, from basic identity recognition to historical periodization and architectural analysis.