Efficient training of language models to fill in the middle
Related stories
Foundations of Large Language Models
Foundations of Large Language Models is a book that focuses on core concepts of large language models rather than exhaustive coverage of the latest technologies. It is organized into six chapters covering pre‑training, generative models, prompting, alignment, inference, and reasoning. The book targets college students, professionals, and practitioners in NLP and related fields, serving as a reference for anyone interested in large language models.
Evaluating large language models trained on code
Beyond the Final Layer: Intermediate Representations for Better Multilingual Calibration in Large Language Models
The paper investigates multilingual confidence calibration in large language models, revealing that non‑English languages are systematically less well calibrated than English. By analyzing internal representations, the authors find that late‑intermediate layers provide a more reliable confidence signal than the final layer, which is biased by English‑centric training. They propose training‑free methods such as Language‑Aware Confidence Ensemble (LACE) to adaptively select optimal layers per language, aiming to improve global equity and trustworthiness of LLMs.
Understanding and Accelerating the Training of Masked Diffusion Language Models
arXiv:2605. 13026v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling.
Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
Multilingual Large Language Models (LLMs) traditionally rely on a single vocabulary shared by all supported languages, which can lead to uneven compression across them. Moreover, their large embedding...
To Each Language Its Tokenizer: Modular Tokenizers for Efficient Multilingual LLMs
The paper proposes a modular tokenizer framework for multilingual large language models, allowing the creation of language‑specific subtokenizers that match monolingual compression quality. It introduces a pretraining strategy that samples these subtokenizers to limit predictions to relevant vocabularies, enabling efficient training and inference. This approach reduces memory usage and speeds up inference without compromising performance.
Improved Large Language Diffusion Models
arXiv:2606. 25331v1 Announce Type: cross Abstract: Modern large language models are predominantly trained with autoregressive factorization and causal attention.
Block Sparse Matrices for Smaller and Faster Language Models
All Entities are Not Created Equal: Examining the Long Tail for Ultra-Fine Entity Typing
arXiv:2410.17355v4 Announce Type: replace Abstract: Due to their capacity to acquire world knowledge from large corpora, pre-trained language models (PLMs) are extensively used in ultra-fine entity t...
Rebalancing Token Importance in Language Models with TF-IDF Weighted Cross-Entropy Loss
Large language models are typically trained under uniform token weighting, which allows frequent and low-information tokens to dominate learning and can increase the tendency to memorize surface-level...