Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.16006v1 Announce Type: cross Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test wha...
The paper introduces the Middle East Cultural Sensitivity Score (MECSS) to quantify Orientalist bias in large language models, converting Said’s seven Orientalist operations into measurable dimensions. Using 280 conversations, it finds that GPT‑4 and Falcon3‑7B‑Instruct systematically reproduce Orientalist patterns, with Falcon scoring higher despite being regionally built. The study highlights that geographic origin alone does not mitigate bias and identifies a new failure mode, "Said‑washing," present in 87.9% of GPT‑4 interactions.
The paper presents an expert-driven method for turning normative principles—specifically Islamic ethical, theological, and jurisprudential traditions—into alignment data for language models. Over a year, seven experts curated 2.8K supervised fine-tuning examples and 5.4K preference pairs in Arabic-English, then evaluated models trained on these datasets. Experiments show that models trained with the curated SFT data outperform a baseline in expert judgments, while adding preference data yields a smaller, non-significant improvement.
The study audits the cultural values expressed by the decision‑only language model JEV using the 2013 Values Survey Module. By presenting 24 items to JEV under 12 matched Saudi and 12 matched American personas, in both English and Arabic, and across eight request formulations, the researchers found that JEV’s responses were highly repeatable (ICC 0.997) and that persona and language significantly influenced the model’s value profiles. Saudi personas shifted JEV’s answers toward the human Saudi‑US difference—capturing 87 % of the effect in English and 62 % in Arabic—while language, age, and gender also modulated the outcomes.
The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.
The study examines how large language models (ChatGPT, Gemini, and Grok) embed religious bias in AI‑generated financial advice. Using 432 simulated advisor‑client interactions across four religious identities and three financial decisions, the authors find that only 12‑18% of advice is unbiased, with Gemini showing the most bias and ChatGPT comparable to Grok. The research identifies structural biases in model design and discursive mechanisms—such as religious anchoring and tone modulation—that vary by scenario, revealing a tension between personalization and neutrality in AI advisory services.