arXiv AI

Right Frame, Wrong Rule: Cultural Cues Expose the Financial Knowledge Gap They Were Meant to Close

arXiv AI
Aug 20

Computational Orientalism: Measuring Structural Discourse Bias in Large Language Models Using the Middle East Cultural Sensitivity Score (MECSS)

The paper introduces the Middle East Cultural Sensitivity Score (MECSS) to quantify Orientalist bias in large language models, converting Said’s seven Orientalist operations into measurable dimensions. Using 280 conversations, it finds that GPT‑4 and Falcon3‑7B‑Instruct systematically reproduce Orientalist patterns, with Falcon scoring higher despite being regionally built. The study highlights that geographic origin alone does not mitigate bias and identifies a new failure mode, "Said‑washing," present in 87.9% of GPT‑4 interactions.

By Maha Shahid
Hugging Face Trending Papers
5d ago

From Normative Frameworks to Alignment Data: Constructing and Evaluating SFT and Preference Data

The paper presents an expert-driven method for turning normative principles—specifically Islamic ethical, theological, and jurisprudential traditions—into alignment data for language models. Over a year, seven experts curated 2.8K supervised fine-tuning examples and 5.4K preference pairs in Arabic-English, then evaluated models trained on these datasets. Experiments show that models trained with the curated SFT data outperform a baseline in expert judgments, while adding preference data yields a smaller, non-significant improvement.

arXiv AI
4d ago

Calibrated to Whom? Persona and Language Effects on Cultural Values in JEV

The study audits the cultural values expressed by the decision‑only language model JEV using the 2013 Values Survey Module. By presenting 24 items to JEV under 12 matched Saudi and 12 matched American personas, in both English and Arabic, and across eight request formulations, the researchers found that JEV’s responses were highly repeatable (ICC 0.997) and that persona and language significantly influenced the model’s value profiles. Saudi personas shifted JEV’s answers toward the human Saudi‑US difference—capturing 87 % of the effect in English and 62 % in Arabic—while language, age, and gender also modulated the outcomes.

By Bushra Asseri, Abdulaziz Asseri
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv AI
Aug 19

When Personalization Becomes Bias: Structural and Discursive Religious Framing in AI-Generated Financial Advice

The study examines how large language models (ChatGPT, Gemini, and Grok) embed religious bias in AI‑generated financial advice. Using 432 simulated advisor‑client interactions across four religious identities and three financial decisions, the authors find that only 12‑18% of advice is unbiased, with Gemini showing the most bias and ChatGPT comparable to Grok. The research identifies structural biases in model design and discursive mechanisms—such as religious anchoring and tone modulation—that vary by scenario, revealing a tension between personalization and neutrality in AI advisory services.

By Muhammad Salar Khan, Hamza Umer, Hasan Mahmud, Sandra Rothenberg
arXiv AI
Jul 23

Overview of FinMMEval 2026 Task 1: Multilingual Financial Multiple-Choice Question Answering

arXiv:2607. 19856v1 Announce Type: cross Abstract: FinMMEval 2026 Task 1 evaluates multilingual financial multiple-choice question answering in English, Chinese, Arabic, and Hindi.

By Zhuohan Xie, Yuyang Dai, Rania Elbadry, Vanshikaa Jani, Georgi Georgiev, Dimitar Dimitrov, Fan Zhang, Xueqing Peng, Lingfei Qian, Jimin Huang, Jiahui Geng, Yankai Chen, Ye Yuan, Haolun Wu, Yuxia Wang, Ivan Koychev, Veselin Stoyanov, Mingzi Song, Yu Chen, Xue Liu, Preslav Nakov
arXiv AI
Sep 10

IGT @ FinMMEval 2026 Task 2: Question-Type Prompting with Targeted Extraction for Multilingual Financial QA

The IGT system tackles PolyFiQA Task 2 of the FinMMEval Lab, a multilingual financial QA challenge involving English SEC filings and news in five languages. It distinguishes two question families: numeric‑structured queries are answered via keyword extraction from filings, while synthesis queries use rule‑based passage selection from news. The approach yields a development ROUGE‑1 of ~0.395, a 60% boost over a generic RAG baseline, and places third among twelve teams on the official test set.

By Yuwen Chiu (Georgia Institute of Technology)
arXiv Computation and Language
Sep 2

Camellia: Benchmarking Cultural Biases in LLMs for Asian Languages

Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.

By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu