OpenAI Blog

Filling crucial language learning gaps

GPT-4 deepens the conversation on Duolingo.

arXiv Computation and Language
6d ago

Feeding BabyLMs Macaroni: Code-Switching Curricula Cause Cross-Lingual Convergence

The paper investigates whether training small decoder-only transformers on code‑switched text can induce cross‑lingual alignment. Using two 100‑million‑word multilingual corpora—one a mix of English, Dutch, and Chinese BabyBabelLM data, and another generated by inserting word‑ and sentence‑level code‑switching via an LLM—the authors find that code‑switched training aligns representations of parallel text, especially across different scripts, and that this alignment persists when later training on monolingual documents. A curriculum that progresses from word‑level code‑switching to sentence‑level code‑switching and finally to monolingual data yields models that outperform baselines on the BabyLM evaluation suite, demonstrating that code‑switching curriculum learning is an effective data augmentation strategy for multilingual pretraining.

By Dries Rooryck, Alex Cai, Yonatan Belinkov, David Alvarez-Melis, Kiant\'e Brantley