arXiv AI By Antoni Czolgowski, Abel Iyasele

Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

Read the original on arXiv AI →

The paper evaluates three open‑weight large language models—Gemma3‑12B (USA), Bielik‑11B‑v3 (Poland), and Qwen3‑4B (China)—against World Values Survey data for 63 demographic personas across three countries, using normalized Wasserstein distance to measure cultural misalignment. Surprisingly, none of the models shows a preference for its home country; Qwen3‑4B, built in China, has the highest misalignment for Chinese respondents. Targeted LoRA fine‑tuning on the five worst‑case personas, with fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik‑11B, but the fine‑tuning redistributes bias rather than eliminating it, shifting worst‑case personas from American to Chinese elderly.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 3

Probing Cultural Signals in Large Language Models through Author Profiling

The study investigates cultural biases in large language models (LLMs) by testing their ability to perform author profiling—inferring singers’ gender and ethnicity—from song lyrics in a zero‑shot setting. Evaluating over 10,000 lyrics across several open‑source models, the authors find that most LLMs default toward North American ethnicity, while DeepSeek‑1.5B leans toward Asian ethnicity, and that Ministral‑8B exhibits the strongest ethnicity bias whereas Gemma‑12B is the most balanced. The paper introduces two fairness metrics, Modality Accuracy Divergence (MAD) and Recall Divergence (RD), to quantify these disparities and provides code and results publicly on GitHub and HuggingFace.

By Valentin Lafargue, Ariel Guerra-Adames, Emmanuelle Claeys, Elouan Vuichard, Jean-Michel Loubes
arXiv AI
Aug 21

DiverValue-Bench: A Benchmark and Fine-Tuning Framework for Aligning Large Language Models with Diverse Human Values

arXiv:2509. 08022v3 Announce Type: replace-cross Abstract: Aligning large language models (LLMs) with diverse human values is essential for safe and effective deployment, yet existing benchmarks often overlook cultural and demographic variation.

By Yao Liang, Dongcheng Zhao, Feifei Zhao, Guobin Shen, Yuwei Wang, Dongqi Liang, Yi Zeng
arXiv AI
Sep 25

Cultural Divergence Preservation: Diagnosing Flattening and Caricature in LLM-Simulated Survey Populations

The paper introduces Cultural Divergence Preservation (CDP), a new diagnostic for evaluating whether large language models (LLMs) preserve cross‑country differences when used as synthetic survey respondents. CDP uses a single human calibration to detect cultural flattening (reduced divergence) or caricature (increased divergence) and is shown to vary monotonically with cross‑country divergence, unlike conventional Jensen–Shannon divergence metrics. Experiments across multiple LLM backbones, prompting methods, and survey domains reveal that CDP uncovers systematic discrepancies with traditional fidelity metrics, highlighting that methods favored by those metrics can still produce strong flattening.

By Yeeun Chae, Yewon Choi, Seunghyun Lee, IL Im