OpenAI Blog
Understanding the capabilities, limitations, and societal impact of large language models
Read the original on OpenAI Blog →The Flow has not summarised this story yet — read it at OpenAI Blog.
The Flow has not summarised this story yet — read it at OpenAI Blog.
The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.