Hugging Face Trending Papers

When Trivia Is Not Trivial: Everyday Knowledge Failures in Multilingual LLMs

Quiz rooms, trivia nights, and quiz shows challenge human knowledge across a wide range of topics, from canonical facts to everyday culture. In this paper, we examine whether large language models (LLMs) can perform competitively in such settings, using quiz-style questions to test them on both common and niche topics.

arXiv AI
Jun 8

The Masked Advantage: Uncovering Local-Language Access to Cultural Knowledge in LLMs

arXiv:2606. 07422v1 Announce Type: cross Abstract: Large language models are increasingly used to answer culturally grounded questions across languages, yet it remains unclear whether local cultural knowledge is better accessed through English or the local language.

By Yang Zhang, Xiao Fei, Amr Mohamed, Sarah Almeida Carneiro, Mersin Konomi, Mingmeng Geng, Ahmed Asaad, Guokan Shang, Michalis Vazirgiannis
arXiv AI
Sep 17

Knowledge-Graph Based Augmentation versus Retrieval Augmented Generation for Cultural-Related Question Answering

The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.

By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv Computation and Language
Sep 25

Language Specific Knowledge: Do Models Know Better in X than in English?

The paper introduces the concept of Language Specific Knowledge (LSK), showing that multilingual language models can answer certain queries better when prompted in a language other than English, sometimes even in low‑resource languages. It defines a language‑selection problem and presents several baseline methods, including the authors’ LSKExtractor, to empirically demonstrate that choosing the optimal language can improve question‑answering performance across datasets covering cultural and social norms. Experiments reveal non‑intuitive mappings, such as Gemma models excelling on Chinese and Middle Eastern topics in Spanish and Qwen models performing best on authority and responsibility queries in Arabic and Chinese.

By Ishika Agarwal, Nimet Beyza Bozdag, Dilek Hakkani-T\"ur
arXiv Machine Learning
Aug 18

L3Cube-IndicQuest v2: A Large-Scale Multilingual Benchmark for Evaluating Factual Knowledge of Large Language Models Across Indic Languages

arXiv:2608. 15535v1 Announce Type: cross Abstract: We present L3Cube-IndicQuest v2, a large-scale gold-standard multilingual question-answering benchmark for evaluating the India-specific factual knowledge of Large Language Models (LLMs).

By Rinit Jain, Tirthraj Mahajan, Advait Joshi, Raviraj Joshi
arXiv AI
Jun 10

RankLLM: Weighted Ranking of LLMs by Quantifying Question Difficulty

arXiv:2602. 12424v2 Announce Type: replace-cross Abstract: Benchmarks establish a standardized evaluation framework to systematically assess the performance of large language models (LLMs), facilitating objective comparisons and driving advancements in the field.

By Ziqian Zhang, Xingjian Hu, Yue Huang, Kai Zhang, Ruoxi Chen, Yixin Liu, Qingsong Wen, Kaidi Xu, Xiangliang Zhang, Neil Zhenqiang Gong, Lichao Sun
Hugging Face Trending Papers
Jun 1

CARTE: A Benchmark for Mapping Language Model Knowledge Across France

We introduce CARTE 1 (Culturally Anchored Regional-Territorial Evaluation), a multiplechoice benchmark for evaluating the ability of large language models (LLMs) to perform fine-grained reasoning over geographically grounded and regionally differentiated knowledge within France. While prior benchmarks focus on national-level cultural understanding, they largely overlook intra-country variation and the need to distinguish between closely related regional contexts.

Hugging Face Trending Papers
Jul 22

D2VBench: Benchmarking Large Language Models with Value Dilemmas in Daily Scenarios

With the wide application of large language models (LLMs) in real-world scenarios, the value implication of their outputs is crucial. However, existing evaluation benchmarks suffer from insufficient coverage of value dilemmas in daily scenarios involving multiple value conflicts and simplistic evaluation formalisms that fail to assess LLMs' value alignment.