How to train a new language model from scratch using Transformers and Tokenizers
Related stories
Beyond the Rosetta Stone: Unification Forces in Generalization Dynamics
The paper investigates why large language models sometimes hallucinate when asked about facts in a language different from the one in which the facts were learned. By training small Transformer models on synthetic multilingual datasets, the authors show that the degree of correlation between facts and their learning language (informativeness) and the ease of language identification (extractability) determine whether models develop unified or separate representations across languages. Unified representations enable cross‑lingual fact transfer, while separate representations do not. The study proposes a unifying perspective on cross‑lingual transfer and suggests training methods to promote representational unification.
TokSuite: Measuring the Impact of Tokenizer Choice on Language Model Behavior
arXiv:2512. 20757v2 Announce Type: replace-cross Abstract: Tokenizers provide the fundamental basis through which text is represented and processed by language models (LMs).
Breaking the Tokenizer Barrier: On-Policy Distillation across Model Families
arXiv:2606. 09456v1 Announce Type: new Abstract: On-Policy Distillation (OPD) has become a core technique in the post-training of Large Language Models (LLMs) for transferring knowledge from domain experts to student models.
How to generate text: using different decoding methods for language generation with Transformers
Tokenization in Transformers v5: Simpler, Clearer, and More Modular
Train and Fine-Tune Sentence Transformers Models
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
The paper proposes a modular tokenizer framework for multilingual large language models, allowing the creation of language‑specific subtokenizers that match monolingual compression quality. It introduces a pretraining strategy that samples these subtokenizers to limit predictions to relevant vocabularies, enabling efficient training and inference. This approach reduces memory usage and speeds up inference without compromising performance.
The Reformer - Pushing the limits of language modeling
How to train a Language Model with Megatron-LM
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, their large embedding...
Distilling Sequential Computation in Transformer Language Models
Transformer language models process sequences token by token in an autoregressive manner, making growing contexts increasingly expensive. Yet many adjacent token spans are highly predictable or freque...