arXiv AI By Haeun Yu, Arnav Arora Seogyeong Jeong, Nadav Borenstein, Siddhesh Pawar, Jisu Shin, Jiho Jin, Junho Myung, Alice Oh, Isabelle Augenstein

CulTrace: Tracing Internal Cultural Reasoning in Large Language Models

Read the original on arXiv AI →

arXiv:2508. 08879v3 Announce Type: replace-cross Abstract: The growing deployment of large language models (LLMs) across diverse cultural contexts necessitates a deeper understanding of models' hidden representations of different cultures.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Jun 8

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.

By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis
arXiv Computation and Language
Aug 31

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.

By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv AI
Sep 2

CHARM: Character Hallucination for Multicultural Role Play Benchmark

CHARM is a multicultural benchmark that tests large language models’ ability to adopt a character’s style while respecting knowledge boundaries. It includes 40 real and fictional characters from five cultural-linguistic regions and evaluates two boundary types—Temporal and Cross-Universe—using abstention-enabled multiple-choice questions. The study finds that hallucinations mainly stem from compliance failures: models often recognize a query is out of scope yet still provide out-of-character answers, revealing systematic cultural variations in these errors.

By Sunkyung Han, Nahyeon Park, Gaeun Seo, Seunghyun Yoon, JinYeong Bak