arXiv AI By Bushra Asseri, Abdulaziz Asseri

Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

Read the original on arXiv AI →

The study audits the cultural values expressed by the decision‑only language model JEV using the 2013 Values Survey Module. By presenting 24 items to JEV under 12 matched Saudi and 12 matched American personas, in both English and Arabic, and across eight request formulations, the researchers found that JEV’s responses were highly repeatable (ICC 0.997) and that persona and language significantly influenced the model’s value profiles. Saudi personas shifted JEV’s answers toward the human Saudi‑US difference—capturing 87 % of the effect in English and 62 % in Arabic—while language, age, and gender also modulated the outcomes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 4

Culturally Grounded Personas in Large Language Models: Characterization and Alignment with Socio-Psychological Value Frameworks

arXiv:2601. 22396v2 Announce Type: replace-cross Abstract: Despite the growing utility of Large Language Models (LLMs) for simulating human behavior, the extent to which these synthetic personas accurately reflect world and moral value systems across different cultural conditionings remains uncertain.

By Candida M. Greco, Lucio La Cava, Andrea Tagarelli
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv AI
Aug 20

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

The paper introduces the Middle East Cultural Sensitivity Score (MECSS) to quantify Orientalist bias in large language models, converting Said’s seven Orientalist operations into measurable dimensions. Using 280 conversations, it finds that GPT‑4 and Falcon3‑7B‑Instruct systematically reproduce Orientalist patterns, with Falcon scoring higher despite being regionally built. The study highlights that geographic origin alone does not mitigate bias and identifies a new failure mode, "Said‑washing," present in 87.9% of GPT‑4 interactions.

By Maha Shahid