Hugging Face Blog

A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality

arXiv AI
Sep 1

Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture

arXiv:2608.28641v1 Announce Type: cross Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Termin...

By Yunsu Kim, Kaden Uhlig, Ashwin Purohit, Milind Agarwal, Patrick Simianer, Anil Arslan, Kiarash Mokhtari, Thomas Zenkel, Johannes Mosig, Gabriel Bretschner, Shamik Bose, Joern Wuebker, John DeNero
arXiv AI
Aug 28

Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics

The paper investigates why large language models sometimes hallucinate when asked about facts in a language different from the one in which the facts were learned. By training small Transformer models on synthetic multilingual datasets, the authors show that the degree of correlation between facts and their learning language (informativeness) and the ease of language identification (extractability) determine whether models develop unified or separate representations across languages. Unified representations enable cross‑lingual fact transfer, while separate representations do not. The study proposes a unifying perspective on cross‑lingual transfer and suggests training methods to promote representational unification.

By Carter Blum, Katja Filippova, Ann Yuan, Asma Ghandeharioun, Julian Zimmert, Fred Zhang, Jessica Hoffmann, Tal Linzen, Martin Wattenberg, Lucas Dixon, Mor Geva
arXiv AI
Sep 4

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.

By Baban Gain, Trilok Nath Singh, Asif Ekbal
arXiv AI
Aug 25

The Multilingual FrameNet Corpus

The paper presents the Multilingual FrameNet Corpus (mFNC), a resource that expands the English Berkeley FrameNet by integrating and harmonizing language‑specific corpora in nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian, and Swedish. Experiments with various model architectures on mFNC consistently surpass existing state‑of‑the‑art Frame Semantic Parsers in both multilingual and cross‑lingual scenarios, highlighting the value of multilingual training data. The mFNC and the trained Frame Semantic Parser models are publicly released on GitHub.

By Beatrice Fiuman\`o, Nicolas Lazzari, Simone Paolo Ponzetto, Valentina Presutti
arXiv Computation and Language
Sep 28

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

The paper investigates whether training small decoder-only transformers on code‑switched text can induce cross‑lingual alignment. Using two 100‑million‑word multilingual corpora—one a mix of English, Dutch, and Chinese BabyBabelLM data, and another generated by inserting word‑ and sentence‑level code‑switching via an LLM—the authors find that code‑switched training aligns representations of parallel text, especially across different scripts, and that this alignment persists when later training on monolingual documents. A curriculum that progresses from word‑level code‑switching to sentence‑level code‑switching and finally to monolingual data yields models that outperform baselines on the BabyLM evaluation suite, demonstrating that code‑switching curriculum learning is an effective data augmentation strategy for multilingual pretraining.

By Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley