The paper "Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities" introduces the Waldo benchmark, a multilingual QA dataset built from Wikipedia that focuses on knowledge gaps and conflicts across languages. It evaluates eight models in five languages and finds that when a fact is missing in one language, models tend to use evidence from the other language, but when conflicting accounts exist, responses align strongly with the query language, leading to different answers for semantically identical questions. The study also explores mitigation strategies, including ablating attention heads and LoRA-based training, which can reduce the preference gap by up to 61.5%.
By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
arXiv:2607. 07707v1 Announce Type: cross Abstract: Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights.
By Yair Feldman, Linxi Zhao, Nathan Godey, Dongyoung Go, Yilun Hua, Kilian Q. Weinberger, Jennifer J. Sun, Yoav Artzi
The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.
By Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur
arXiv:2610.06650v2 Announce Type: replace
Abstract: Wikidata is one of the largest open knowledge bases, yet answering a complex question over it still requires a SPARQL query that names the right en...
By Mohamed Chenene, Carlos Rosas-Hinostroza, Anastasia Stasenko, Shani Evenstein Sigalov, Pierre-Carl Langlais
arXiv:2607. 06327v1 Announce Type: cross Abstract: Uncertainty estimation (UE) enables LLM-powered systems to recognize when to abstain, yet existing research has predominantly focused on English.
By Andrea Alfarano, Andrea Bacciu, Saab Mansour, Amin Mantrach, Marcello Federico
The paper introduces Knowledge-Weighted Fine‑Tuning, a method that estimates an instance‑level knowledge score through multi‑sampled inference and uses it to scale the learning signal. This approach encourages large language models to explicitly say "I don't know" on out‑of‑scope queries while preserving accuracy on known questions. The authors also propose new evaluation metrics for uncertainty, demonstrating that better discrimination between known and unknown instances improves overall performance.
By Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim
arXiv:2510.13935v3 Announce Type: replace-cross
Abstract: The facts a language model stores are tied to its parameter count, so small models that fit on edge devices fail on expert problems, which ne...
By Kenan Alkiek, David Jurgens, Vinod Vydiswaran
arXiv:2604. 23336v3 Announce Type: replace-cross Abstract: Unlike traditional fact-based retrieval, rationale-based retrieval typically necessitates cross-encoding of query-document pairs using large language models, incurring substantial computational costs.
By Teng Chen, Sheng Xu, Feixiang Guo, Xiaoyu Wang, Qingqing Gu, Hongyan Li, Luo Ji
arXiv:2502. 15631v2 Announce Type: replace-cross Abstract: Large language models have demonstrated remarkable progress in mathematical reasoning, leveraging chain-of-thought and reinforcement learning.
By Marthe Ballon, Andres Algaba, Vincent Ginis
The study explores how the language used for reasoning affects retrieval‑augmented generation (RAG) in a monolingual German setting. Using a German RAG question‑answering testbed based on the tabletop game The Dark Eye, the authors show that aligning the reasoning language with the query and retrieved documents improves performance, with German reasoning outperforming French reasoning. However, German reasoning still does not surpass the model’s native English reasoning, indicating that native multilingual reasoning is necessary for optimal results.
By Oliver Hauck, Mario Sanz-Guerrero, Katharina von der Wense
arXiv:2606. 19349v1 Announce Type: cross Abstract: While In-Context Learning (ICL) is extensively studied in Autoregressive (AR) LLMs, its mechanism within Diffusion Large Language Models (dLLMs) remains largely unexplored.
By Zhengheng Li, Panrui Li, Xuyang Liu, Puzhi Xia
The paper evaluates how robust large language models (LLMs) are when using retrieval‑augmented generation (RAG) in practical settings. It investigates whether RAG always outperforms non‑RAG approaches, whether adding more retrieved documents helps, and whether the order of documents matters, using a benchmark of 1,891 samples across five datasets and three task categories. Experiments with 11 LLMs show generally high retrieval robustness, but performance varies by task and prompting strategy, indicating that adopting RAG should be considered on a case‑by‑case basis.
By Shuyang Cao, Karthik Radhakrishnan, David Rosenberg, Steven Lu, Pengxiang Cheng, Lu Wang, Shiyue Zhang