A Deepdive into Aya Expanse: Advancing the Frontier of Multilinguality
Related stories
Introducing The World's Largest Open Multilingual Language Model: BLOOM
Llama 3.1 - 405B, 70B & 8B with multilinguality and long context
Terminal-Bench-LILT: Multilingual Agentic Coding Benchmark Grounded in Language, Region, and Culture
arXiv:2608.28641v1 Announce Type: cross Abstract: Most evaluations for coding agents are conducted exclusively in English, which does not reflect real-world multilingual deployment. We present Termin...
On the Limitations of Cross-Lingual Consistency in Multilingual Text-to-image Generation
arXiv:2608. 11002v1 Announce Type: cross Abstract: Text-to-image (T2I) generation has achieved remarkable progress in recent years.
Visual Document Retrieval Goes Multilingual
Granite Embedding Multilingual R2: Open Apache 2.0 Multilingual Embeddings with 32K Context — Best Sub-100M Retrieval Quality
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
The paper investigates why large language models sometimes hallucinate when asked about facts in a language different from the one in which the facts were learned. By training small Transformer models on synthetic multilingual datasets, the authors show that the degree of correlation between facts and their learning language (informativeness) and the ease of language identification (extractability) determine whether models develop unified or separate representations across languages. Unified representations enable cross‑lingual fact transfer, while separate representations do not. The study proposes a unifying perspective on cross‑lingual transfer and suggests training methods to promote representational unification.
No Optimal Language Set Exists for Multilingual Instruction Tuning: Insights from a Linguistically-Informed Study
arXiv:2410. 07809v2 Announce Type: replace-cross Abstract: Multilingual instruction tuning (MIT) is challenged by the curse of multilinguality, data scarcity, and high computational cost.
One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging
The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.
The Multilingual FrameNet Corpus
The paper presents the Multilingual FrameNet Corpus (mFNC), a resource that expands the English Berkeley FrameNet by integrating and harmonizing language‑specific corpora in nine additional languages: Brazilian Portuguese, Chinese, Dutch, French, German, Italian, Korean, Latvian, and Swedish. Experiments with various model architectures on mFNC consistently surpass existing state‑of‑the‑art Frame Semantic Parsers in both multilingual and cross‑lingual scenarios, highlighting the value of multilingual training data. The mFNC and the trained Frame Semantic Parser models are publicly released on GitHub.
Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence
The paper investigates whether training small decoder-only transformers on code‑switched text can induce cross‑lingual alignment. Using two 100‑million‑word multilingual corpora—one a mix of English, Dutch, and Chinese BabyBabelLM data, and another generated by inserting word‑ and sentence‑level code‑switching via an LLM—the authors find that code‑switched training aligns representations of parallel text, especially across different scripts, and that this alignment persists when later training on monolingual documents. A curriculum that progresses from word‑level code‑switching to sentence‑level code‑switching and finally to monolingual data yields models that outperform baselines on the BabyLM evaluation suite, demonstrating that code‑switching curriculum learning is an effective data augmentation strategy for multilingual pretraining.