arXiv AI By Christelle Clervilsson, Yanzhu Guo

Evaluating Cultural Awareness of LLMs for Haitian Creole

Read the original on arXiv AI →

The paper presents the first systematic evaluation of cultural awareness in large language models (LLMs) for Haitian Creole, a low‑resource language. Using a benchmark of culturally salient prompts curated by native speakers, the study assesses four dimensions—specificity, bias, diversity, and variation—in a text‑infilling setting. Results show a clear gap between Haitian Creole and higher‑resource French, with Haitian performance more uneven and more affected by French linguistic interference; story generation reveals stereotypical portrayals of Haitian characters even in positive contexts.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 25

Beyond Surface Cues: Disentangling Sociocultural Signals in Multilingual LLMs

The study examines how multilingual large language models (LLMs) produce outputs that differ across sociocultural contexts, highlighting that identity labels and source-language cues can mislead assessments of cultural grounding. Using a human‑validated, multi‑agent audit on 89,253 outputs from 12 LLMs in English, French, and Chinese across 18 occupations and three task conditions, the authors find that bias representation varies systematically by language and task. Removing direct identity cues reduces identity‑label prediction in English and Chinese but not in French, and the source language’s cultural context consistently receives the highest relevance score, though this signal weakens after translation or name masking. "whyItMatters":"The findings show that surface cues can obscure true cross‑cultural patterns, underscoring the need for careful audit designs to avoid misleading conclusions about bias in multilingual LLMs."

By Yuanjun Feng, Tanzhou Liu, Stefan Feuerriegel, Yash Raj Shrestha
Hugging Face Trending Papers
Jun 1

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts.

arXiv Machine Learning
Aug 5

VIVID: A Culturally Grounded Benchmark Exposing the Figurative Language Gap in Vietnamese NLP

arXiv:2608. 03095v1 Announce Type: cross Abstract: We present VIVID (Vietnamese Idioms for Validation and Interpretation Depth), the first systematic benchmark for evaluating culturally grounded figurative language understanding in Vietnamese.

By Tu Tran Do, Nhat Ngoc Nguyen, Khanh-Tung Tran, Hoang D. Nguyen, Tu Minh Phuong, Long Hoang Dang
arXiv AI
Jul 3

Challenges and Recommendations for LLMs-as-a-Judge in Multilingual Settings and Low-Resource Languages

arXiv:2607. 02235v1 Announce Type: cross Abstract: LLM-as-a-Judge has become the dominant evaluation paradigm for many natural language generation tasks, due to shortcomings of conventional metrics and high correlations with human judgment, albeit mostly in English.

By A. Seza Do\u{g}ru\"oz, Xixian Liao, Verena Blaschke, Jakob Prange, Senyu Li, David Ifeoluwa Adelani
arXiv AI
Sep 21

Lessons Without Borders? Evaluating Cultural Alignment of LLMs Using Multilingual Story Moral Generation

The paper introduces a multilingual story moral generation task to evaluate cultural alignment in large language models. Using a dataset of human-written story morals from 14 language‑culture pairs, the authors compare model outputs to human interpretations through semantic similarity, a preference survey, and value categorization. They find that advanced models like GPT‑4o and Gemini produce morally similar and preferred responses but show less cross‑linguistic variation, focusing on a narrower set of shared values, indicating a limitation in capturing the diversity of human narrative understanding.

By Sophie Wu, Andrew Piper