arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
CultureVidBench is a new benchmark that evaluates how well text‑to‑video generation models capture cultural details. It contains 1,000 prompts spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects, grouped into material culture, social practice & performance, and ritual & ceremony. Human studies and automated assessments show that while current models perform well on semantic adherence and visual quality, they often miss fine‑grained cultural details, especially for underrepresented regions and multimodal cues.
By Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
TempCloze is a video cloze benchmark designed to evaluate visual temporal reasoning in Video-LLMs. The task presents a video’s beginning and ending clips and asks models to select the correct missing middle from four candidates, focusing on semantic, alignment, and progression aspects while minimizing appearance cues. Evaluation of 31 models shows that temporal alignment is the main challenge, with models performing better on semantic content and event progression but struggling to place events correctly in time.
By Wenqi Pei, Henry Hengyuan Zhao, Yilai Liu, Jiahao Meng, Han Chen, Ziyu Wang, Hongyang Du
arXiv:2609.01772v1 Announce Type: new
Abstract: Meme understanding goes beyond recognizing visual content or literal text; it requires implicit cultural knowledge and pragmatic inference that most vi...
By Tawsif Tashwar Dipto, Mehedi Ahamed, Radib Bin Kabir, Mueeze Al Mushabbir, Mohammed Saidul Islam, Mir Rayat Imtiaz Hossain, Md Tahmid Rahman Laskar, Sabbir Ahmed
arXiv:2606. 07311v1 Announce Type: cross Abstract: As video generation models like Veo 3.
By Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes, Mahdi M. Kalayeh, Paul Pu Liang
arXiv:2605. 16716v5 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
By Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat