Hugging Face Blog

Introducing The World's Largest Open Multilingual Language Model: BLOOM

arXiv AI
Sep 2

The Curse of Multilinguality in Lexical Normalization

The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.

By Saman Rahbar
arXiv Computation and Language
Aug 28

Cross-lingual Representation Learning via Centroid Intervention Fusion

The paper introduces Centroid Intervention Fusion (CIF), a framework that merges multiple multilingual intervention projections into a single language-shared operator for inference-time modification of large language models. CIF improves cross-lingual transfer without updating model parameters and achieves up to +3.378 percentage points better performance than prior pairwise intervention baselines across several benchmarks, including low-resource languages. The authors provide code at https://github.com/VRCMF/CIF.git.

By Wei Sun, Marie-Francine Moens
arXiv AI
Sep 4

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.

By Baban Gain, Trilok Nath Singh, Asif Ekbal
arXiv AI
Aug 28

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

The paper investigates why large language models sometimes hallucinate when asked about facts in a language different from the one in which the facts were learned. By training small Transformer models on synthetic multilingual datasets, the authors show that the degree of correlation between facts and their learning language (informativeness) and the ease of language identification (extractability) determine whether models develop unified or separate representations across languages. Unified representations enable cross‑lingual fact transfer, while separate representations do not. The study proposes a unifying perspective on cross‑lingual transfer and suggests training methods to promote representational unification.

By Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva