arXiv Computation and Language

Language Specific Knowledge: Do Models Know Better in X than in English?

The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.

arXiv AI
Jun 8

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.

By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis
arXiv AI
Sep 7

A Systematic Evaluation of Cross-Lingual Consistency Enhancement Methods in Multilingual Language Models

The paper presents a unified evaluation of cross‑lingual consistency (CLC) enhancement methods for multilingual language models, covering inference‑time interventions and post‑training approaches across three model families and three closed‑form benchmarks. Results indicate that post‑training methods, especially direct distribution alignment, consistently improve CLC across all model‑dataset combinations, while other methods are more sensitive to answer format and language coverage. The study also examines the impact of CLC enhancement on culturally diverse question answering, finding no systematic degradation in controlled settings but occasional accuracy drops in open‑ended generation, particularly for non‑English responses.

By Jirui Qi, Mingyang Wang, Hinrich Sch\"utze, Raquel Fern\'andez, Arianna Bisazza
arXiv AI
Sep 17

Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.

By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv Computation and Language
Aug 31

Why are all LLMs Obsessed with Japanese Culture? On the Hidden Cultural and Regional Biases of LLMs

The paper investigates cultural biases in large language models (LLMs) by introducing the Culture-Related Open Questions (CROQ) dataset, which contains 24‑language questions about generic culture. Experiments reveal that LLMs disproportionately favor Japan in their responses, especially when prompted in high‑resource languages, while low‑resource languages tend to highlight countries where the language is official. The study also finds that these biases emerge after supervised fine‑tuning rather than during pre‑training.

By Joseba Fernandez de Landa, Carla Perez-Almendros, Jose Camacho-Collados
arXiv Computation and Language
Sep 17

M-SQE: Multilingual Skill Quality Estimation for Enhancing Language Equality in Agentic Skill Use

M‑SQE is a post‑retrieval framework that estimates the quality of multilingual agent skills by combining a Theory view (intrinsic quality) and an Action view (task‑grounded utility) into a domain‑conditioned score. It was evaluated on general, tool‑use, and cultural skill‑use domains, showing a task‑success improvement of at least +3.5 points over baselines across three retrievers. The method notably boosts performance for low‑resource languages, raising Hindi by +12.9 pp and Swahili by +5.6 pp, and achieves strong results across six cultural regions, advancing linguistic and cultural equality in agentic skill use.

By Yilun Liu, Shimin Tao, Minggui He, Chenxin Liu, Li Zhang, Chen Liu, Miao Zhang, Jiaxin Guo, Min Zhang, Liqun Deng, Xiaojun Meng, Daimeng Wei
arXiv AI
Sep 1

Do Language Models Reason Across Languages?

The paper investigates whether language models can reason across languages by introducing a two‑hop question answering task that requires inference over two multilingual documents. Results show that models are more sensitive to language variation in answer‑span documents than in bridging documents, and that up to 33% of multilingual cases involve correct final answers despite failing to infer bridging information in the first step. The study also reveals an 18% composition failure rate and proposes a three‑stage SUBQ prompting method that improves accuracy from 10.1% to 66.5%.

By Yan Meng, Wafaa Mohammed, Christof Monz
arXiv Computation and Language
4d ago

Large Language Model Selection with Limited Annotations

arXiv:2605.24981v2 Announce Type: replace Abstract: Choosing a Large Language Model (LLM) for a given task requires comparing many strong candidates, yet standard evaluation relies on costly annotati...

By Yavuz Durmazkeser, Patrik Okanovic, Andreas Kirsch, Torsten Hoefler, Nezihe Merve G\"urel