arXiv AI By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

Read the original on arXiv AI →

arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Sep 25

Language Specific Knowledge: Do Models Know Better in X than in English?

The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.

By Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur
arXiv Computation and Language
Aug 31

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.

By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv Computation and Language
Aug 28

MAPLE: Metadata Conditioned LLM Pretraining for Locale-Aware Question Answering

The paper introduces MAPLE, a family of decoder‑only language models pretrained with document‑level geographic metadata such as source URL, country, and continent. MAPLE is evaluated on a new benchmark, LocalNewsQA, which tests whether models can switch answers when the locale changes. Experiments show that, with inference‑time metadata fixed, MAPLE outperforms metadata‑free controls in both answer switching and accuracy on locale‑dependent questions, and these gains grow with model size.

By Anjishnu Mukherjee, Ziwei Zhu, Antonios Anastasopoulos