arXiv AI

Cultural Binding Heads in Language Models

arXiv:2605. 28543v2 Announce Type: replace Abstract: LLMs often default to equal treatment across cultural groups, even though context warrants differentiation: this is a lack of difference awareness.

arXiv AI
Aug 18

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

arXiv:2508. 08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures.

By Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein
arXiv AI
Sep 21

Decoupling Internal Representational Changes and Causal Importance in Fine-Tuned Large Language Models

Fine‑tuning reshapes internal representations of large language models, affecting attention patterns and layer‑wise activations. The study shows that components identified by EAP as important for task performance cluster in specific layers, yet these layers do not align with those undergoing the largest representational changes. Additionally, overlapping EAP components across different tasks do not guarantee cross‑task transfer and can even degrade performance when tasks differ in nature.

By Lingfang Li, Procheta Sen, Shubham Das, Danushka Bollegala
arXiv AI
Sep 7

Cultural Misalignment in Large Language Models: Detection, Measurement, and Mitigation Through Targeted Fine-Tuning

The paper evaluates three open‑weight large language models—Gemma3‑12B (USA), Bielik‑11B‑v3 (Poland), and Qwen3‑4B (China)—against World Values Survey data for 63 demographic personas across three countries, using normalized Wasserstein distance to measure cultural misalignment. Surprisingly, none of the models shows a preference for its home country; Qwen3‑4B, built in China, has the highest misalignment for Chinese respondents. Targeted LoRA fine‑tuning on the five worst‑case personas, with fewer than 1,200 training pairs and under 15 minutes on a single GPU, reduces bias by 16.8% for Bielik‑11B, but the fine‑tuning redistributes bias rather than eliminating it, shifting worst‑case personas from American to Chinese elderly.

By Antoni Czolgowski, Abel Iyasele