The study investigates how different prompt components affect language model responses in psychometric tests. By crossing five distinct baseline personas with five variants of each prompt element—persona wording, task instruction, item wording, and option symbol—the authors measure response shifts using the 1‑Wasserstein distance. Their analysis of 13 small open‑weight language models on the Big Five Inventory and Short Dark Triad reveals that task instruction and option symbol changes often cause more variation than paraphrasing the persona or item, with prompt artifacts explaining over 50% of the variation for many items.
By Nils Schwager, Christoph Hau, Simon M\"unker, Achim Rettinger
Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias. While this behavior has been widely studied for general text generation, its impact on code generation quality and programming conventions remains largely unexplored.
arXiv:2607. 14816v1 Announce Type: cross Abstract: Large Language Models (LLMs) perform differently on identical programming tasks when prompted in different natural languages, a phenomenon known as language bias.
By Saima Afrin, Alessandro Midolo, Camilo Escobar-Vel\'asquez, Mario Linares-V\'asquez, Weiyuan Ding, Bowen Xu, Massimiliano Di Penta, Antonio Mastropaolo
arXiv:2607. 13568v1 Announce Type: cross Abstract: Can a language model estimate its familiarity with an entity before generating an answer?
By Grzegorz Brzezinka
arXiv:2512. 20757v2 Announce Type: replace-cross Abstract: Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs).
By G\"ul Sena Alt{\i}nta\c{s}, Malikeh Ehghaghi, Brian Lester, Fengyuan Liu, Wanru Zhao, Marco Ciccone, Colin Raffel
The paper proposes the interlingua hypothesis, suggesting that large language models translate by encoding a source sentence into a latent, task‑agnostic feature space and then decoding a target sentence from that space. Three lines of evidence support this: (1) BLEU variance across language pairs is largely explained by language‑specific competences without pair‑specific interactions; (2) many model components influence both monolingual and translation tasks; and (3) fine‑tuning on monolingual data recovers most translation gains seen with aligned documents. These findings converge to support the hypothesis and point toward new ways to understand and improve LLM translation.
By Jacob Brinton, Jannik Brinkmann, Mark Crovella, Aaron Mueller