arXiv:2605. 16716v5 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
By Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat
arXiv:2510. 08543v2 Announce Type: replace-cross Abstract: As Video Large Language Models (VideoLLMs) are deployed globally, it is important to assess their ability to reason across cultural contexts.
By Nikhil Reddy Varimalla, Yunfei Xu, Meng Fan Wang, Arkadiy Saakyan, Smaranda Muresan
arXiv:2605. 16716v4 Announce Type: replace-cross Abstract: Text-to-video (T2V) generation has rapidly progressed in visual fidelity, yet its ability to faithfully represent multiple cultures within a single prompt remains underexplored.
By Shuowei Li, Yuming Zhao, Parth Bhalerao, Oana Ignat
Lexara-RF introduces reference‑free metrics for evaluating conversational visual analytics agents that generate visualizations and natural‑language explanations. The framework uses only the prompt, data, and model response to score outputs, applying 13 metrics derived from visualization design theory and Gricean principles as consistency, intent‑alignment, and design validity checks. In tests against a human‑rated corpus, Lexara‑RF matches reference‑based methods, outperforms surface‑similarity NLG baselines, and accurately identifies structurally grounded failures.
By Srishti Palani, Vidya Setlur
CultureVidBench is a new benchmark that evaluates how well text‑to‑video generation models capture cultural details. It contains 1,000 prompts spanning 12 countries, 6 continents, 8 cultural regions, and 14 cultural aspects, grouped into material culture, social practice & performance, and ritual & ceremony. Human studies and automated assessments show that while current models perform well on semantic adherence and visual quality, they often miss fine‑grained cultural details, especially for underrepresented regions and multimodal cues.
By Xianjing Han, Yuhan Su, Yang Deng, Dong Ma, Wee Peng Tay, Bin Zhu
arXiv:2606. 01897v1 Announce Type: new Abstract: Traditional Video Quality Assessment (VQA) focuses narrowly on aesthetic fidelity, overlooking the complex social dynamics that define quality in User-Generated Content (UGC).
By Tianjiao Li, Kai Zhao, Xiang Li, Yang Liu, Huyang Sun
CuBEs introduces culturally‑situated behavioral evaluations for large language models, adding cultural context to test scenarios and assessments. The authors built a human‑labeled dataset covering 12 cultures, revealing significant cross‑cultural differences that one‑size‑fits‑all judgments miss. Evaluating 13 LLMs shows that culturally situated tests uncover varied behaviors, such as Western political bias versus non‑Western religious or colonial biases, which standard evaluations overlook.
By Hoda Ayad, Tanu Mitra, Abhishek Mukherji
The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.
By Sophie Wu, Andrew Piper
arXiv:2606. 13397v1 Announce Type: cross Abstract: Language operates as a mechanism of both marginalization and resistance, especially for minority communities navigating insensitive and harmful speech online.
By Dipto Das, Achhiya Sultana, Ankit Singh Chauhan, Saadia Binte Alam, Mohammad Shidujaman, Shion Guha, Sunandan Chakraborty, Syed Ishtiaque Ahmed
The paper introduces ChartBias, a benchmark of 820 real-world charts covering six social attributes, designed to audit bias in vision‑language models (VLMs) that interpret charts. Across 12 VLMs, the study identifies three failure modes—narrative shift, group hallucination, and preference polarity—where models produce different or misleading narratives when the referenced social group changes. A multi‑agent mitigation framework is proposed, separating evidence extraction from group‑conditioned generation and using a counterfactual judge, which reduces narrative shift while maintaining chart‑grounded reasoning.
By Mizanur Rahman, Huan Wu, Arash Asgari, Enamul Hoque Prince, Laleh Seyyed-Kalantari
arXiv:2606. 00369v1 Announce Type: cross Abstract: Safe global deployment of AI models requires alignment with human values that vary across cultures.
By Arkadiy Saakyan, Charvi Rastogi, Lora Aroyo
arXiv:2606. 07311v1 Announce Type: cross Abstract: As video generation models like Veo 3.
By Anku Rani, Wei Dai, Shravan Nayak, Pattie Maes, Mahdi M. Kalayeh, Paul Pu Liang