arXiv Computation and Language By Lihu Chen

The Geometry of Knowledge Accessibility in Large Language Models

Read the original on arXiv Computation and Language →

The paper investigates how easily large language models (LLMs) can retrieve knowledge for a given query, introducing the concept of knowledge accessibility. It discovers that the accessibility of a query is reflected in its geometric position in the model’s representation space: queries closer to a central point are more accessible, while those farther away are less so. This geometric insight identifies a knowledge boundary, shows that accessibility ordering is consistent across datasets, and informs which interventions—such as query rewriting, chain-of-thought reasoning, or retrieval—are most effective depending on a query’s position relative to the center.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Oct 2

Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities

The paper "Where's Waldo? Query-language Preference under Cross-lingual Knowledge Disparities" introduces the Waldo benchmark, a multilingual QA dataset built from Wikipedia that focuses on knowledge gaps and conflicts across languages. It evaluates eight models in five languages and finds that when a fact is missing in one language, models tend to use evidence from the other language, but when conflicting accounts exist, responses align strongly with the query language, leading to different answers for semantically identical questions. The study also explores mitigation strategies, including ablating attention heads and LoRA-based training, which can reduce the preference gap by up to 61.5%.

By Dayeon Ki, Ruochen Zhang, Silviu Cucerzan, Ryen W. White, Ning Gao
arXiv AI
Jul 9

Co-LMLM: Continuous-Query Limited Memory Language Models

arXiv:2607. 07707v1 Announce Type: cross Abstract: Limited memory language models (LMLMs) externalize factual knowledge during pretraining to a knowledge base (KB), rather than memorizing it in their weights.

By Yair Feldman, Linxi Zhao, Nathan Godey, Dongyoung Go, Yilun Hua, Kilian Q. Weinberger, Jennifer J. Sun, Yoav Artzi
arXiv Computation and Language
Sep 25

Language Specific Knowledge: Do Models Know Better in X than in English?

The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.

By Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur
arXiv AI
Aug 25

What Models Know, How Well They Know It: Knowledge-Weighted Fine-Tuning for Learning When to Say "I Don't Know"

The paper introduces Knowledge-Weighted Fine‑Tuning, a method that estimates an instance‑level knowledge score through multi‑sampled inference and uses it to scale the learning signal. This approach encourages large language models to explicitly say "I don't know" on out‑of‑scope queries while preserving accuracy on known questions. The authors also propose new evaluation metrics for uncertainty, demonstrating that better discrimination between known and unknown instances improves overall performance.

By Joosung Lee, Hwiyeol Jo, Donghyeon Ko, Kyubyung Chae, Cheonbok Park, Jeonghoon Kim