arXiv Computation and Language

Language, Language Models, and What We're Talking About

The paper discusses how language models, often treated as technical artifacts, are actually shaped by the linguistic data used in their training. Using Italian language models trained on translated and synthetic data, the author questions whether these models truly represent Italian or language more broadly, and whether NLP should focus on producing natural language. The work calls for a clearer distinction between models built as products and those built as tools for linguistic study, suggesting that diverse answers and languages may emerge without necessarily being pessimistic.

arXiv AI
Aug 20

Language Models for Portuguese: A Systematic Mapping Study

The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.

By Jhessica Silva, Carlos Caetano, Helena Maia, Breno Bernard Nicolau de Fran\c{c}a, Sandra Avila, Helio Pedrini
arXiv Computation and Language
Aug 25

Dialects of Translationese Shape Language Model Learning

The paper investigates how machine‑translated English data from 24 diverse source languages influences small English language models. It finds that source language affects model behavior: lexical diversity drives overall perplexity, while grammatical performance correlates with typological similarity to English when sufficient data is used. Additionally, translation quality strongly predicts language‑modeling performance.

By Jenny Kunz
arXiv AI
Aug 28

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

The paper investigates why large language models sometimes hallucinate when asked about facts in a language different from the one in which the facts were learned. By training small Transformer models on synthetic multilingual datasets, the authors show that the degree of correlation between facts and their learning language (informativeness) and the ease of language identification (extractability) determine whether models develop unified or separate representations across languages. Unified representations enable cross‑lingual fact transfer, while separate representations do not. The study proposes a unifying perspective on cross‑lingual transfer and suggests training methods to promote representational unification.

By Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva
arXiv AI
Sep 2

The Curse of Multilinguality in Lexical Normalization

The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.

By Saman Rahbar
arXiv Computation and Language
Sep 16

surprisal is Not a Theory

The article argues that Surprisal Theory, often presented as a computational-level explanation, is not a theory in its own right. It contends that using large language model (LLM) surprisals without considering the underlying representational and algorithmic choices obscures the theory’s commitments. The authors demonstrate through three analyses that algorithm and architecture significantly influence language model probabilities, urging researchers to reassess treating LLM surprisals as interchangeable.

By Andr\'es Bux\'o-Lugo, Aniello De Santo, Morgan Grobol, Ryan J. Hubbard, Cassandra L. Jacobs