Hugging Face Trending Papers

Language Equality has a Price: A Systematic Investigation of Multi-turn LLM Performance for EU-24+

Read the original on Hugging Face Trending Papers →

We evaluate large language models (LLMs) as language agents playing goal-directed dialogue games in self-play across 30 languages: the 24 official EU languages plus six others. Unlike static or preference-based evaluation, this paradigm is multi-turn, reference-free and programmatically scored, and because the game mechanics are language-agnostic it extends to a new language by localising a fixed set of prompt and word-list files.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv Computation and Language
Aug 27

Lower-Resource, Higher Scores: Language Bias in LLM Evaluators

The paper demonstrates that large language model (LLM) evaluators, whether reward‑model based or prompted LLM‑as‑a‑Judge, exhibit significant language bias in multilingual settings. Experiments with semantically identical instruction‑response pairs across 23 languages reveal that lower‑resource languages receive higher scores, a bias that persists across eight open‑weight evaluators and is not detectable by standard pairwise accuracy metrics. The authors link the bias to model uncertainty and language identity, showing it cannot be explained by content difficulty alone.

By Ej Zhou, Lucas Resck, Zheng Hui, Anna Korhonen
arXiv Machine Learning
Aug 27

Skill Issue: Are Skills Language-Invariant in LLMs?

The paper investigates whether large language models (LLMs) exhibit language‑specific skill differences by having two identical model instances play a text‑based game in different languages. Using a multilingual extension of TextArena, the authors evaluate three open‑weight models across eight languages and six games, finding that the same model can show markedly different performance—varying win–loss margins, invalid actions, and strategic choices—depending on the language interface. Analyses pinpoint language‑specific failures in spatial reasoning, card‑conditioned decisions, and optimal move selection, and demonstrate that adjusting the intermediate reasoning language can recover much of the lost performance.

By Bobby Cheng, Adam Gaber, Zhengyuan Liu, Catherine Arnett, Omer Goldman, Cheston Tan, Leshem Choshen