arXiv AI By Hoda Ayad, Tanu Mitra, Abhishek Mukherji

CuBEs: Culturally-Situated Behavioral Evaluations and the Limitations of Culture-Blind LLM Judges

Read the original on arXiv AI →

CuBEs introduces culturally‑situated behavioral evaluations for large language models, adding cultural context to test scenarios and assessments. The authors built a human‑labeled dataset covering 12 cultures, revealing significant cross‑cultural differences that one‑size‑fits‑all judgments miss. Evaluating 13 LLMs shows that culturally situated tests uncover varied behaviors, such as Western political bias versus non‑Western religious or colonial biases, which standard evaluations overlook.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 31

CultureConverse: A Multilingual Multi-turn Simulation Harness for Culturally Grounded Assistance in East and Southeast Asia

arXiv:2608.28405v1 Announce Type: new Abstract: Current cultural evaluations for large language models (LLMs) often reduce culture to single-turn factual recall via MCQs, failing to capture a common...

By Bryan Chen Zhengyu Tan, Weihua Zheng, Thong T. Doan, Bich Ngoc Doan, Jia Wang Peh, Xiaoyuan Yi, Jing Yao, Xing Xie, Nancy F. Chen, Zhengyuan Liu, JinYeong Bak, Wafi Shamdi, Soo Kai Chie, Liew Yu Siong, Aina Azyyati Binti Mohamad Rezal, Lew Yan Yan Vanessa, Huadan Wu, Dylan Raharja, Nadya Yuki Wangsajaya, Akane Fukushige, Kazushi Kato, Koji Inoue, Tatsuya Kawahara, Jaehyung Seo, Dongjun Kim, Seungyoon Lee, Zi Haur Pang, Rui Yang Tan, Charibeth Ko Cheng, Maria Regina Justina Estuar, Jann Railey Montalan, Pham Minh Duc, Roy Ka-Wei Lee
arXiv AI
Sep 25

Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations

The paper introduces Cultural Divergence Preservation (CDP), a new diagnostic for evaluating whether large language models (LLMs) preserve cross‑country differences when used as synthetic survey respondents. CDP uses a single human calibration to detect cultural flattening (reduced divergence) or caricature (increased divergence) and is shown to vary monotonically with cross‑country divergence, unlike conventional Jensen–Shannon divergence metrics. Experiments across multiple LLM backbones, prompting methods, and survey domains reveal that CDP uncovers systematic discrepancies with traditional fidelity metrics, highlighting that methods favored by those metrics can still produce strong flattening.

By Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im