Cosmopedia: how to create large-scale synthetic data for pre-training Large Language Models
Related stories
How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data
arXiv:2604. 13977v2 Announce Type: replace-cross Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent.
Very Large Language Models and How to Evaluate Them
Learning Spatio-Temporal Foundation Models from Pure Synthetic Data
arXiv:2607. 16251v1 Announce Type: new Abstract: Spatio-Temporal Foundation Models (STFMs) aim to learn generalizable representations of complex dynamical systems across space and time.
Efficient training of language models to fill in the middle
Towards Engineering Scaling Laws with Pretraining Data Composition
arXiv:2606. 19781v1 Announce Type: cross Abstract: Neural scaling laws describe how model performance improves as a power law in compute, model size, and dataset size.
Speculative Decoding and the Curse of Multilinguality
arXiv:2605. 30580v2 Announce Type: replace-cross Abstract: Speculative decoding is a popular technique for large language model (LLM) inference, enabling faster generation by drafting multiple tokens with a smaller draft model.
FineInstructions: Scaling Synthetic Instructions to Pre-Training Scale
arXiv:2601. 22146v2 Announce Type: replace-cross Abstract: Due to limited supervised training data, large language models (LLMs) are typically pre-trained via a self-supervised "predict the next word" objective on a vast amount of unstructured text data.
Dango: A Strictly L1-Only Large Language Model for Studying Second Language Acquisition
We introduce Dango, a 1. 8B-parameter large language model designed for controlled studies of L1-to-L2 (Japanese-to-English) transfer in second language acquisition (SLA).
Introducing the Synthetic Data Generator - Build Datasets with Natural Language
The State Of LLMs 2025: Progress, Problems, and Predictions
A 2025 review of large language models, from DeepSeek R1 and RLVR to inference-time scaling, benchmarks, architectures, and predictions for 2026.
Overcoming the Modality Gap in Context-Aided Forecasting
arXiv:2603. 12451v4 Announce Type: replace Abstract: Context-aided forecasting (CAF) holds promise for integrating domain knowledge and forward-looking information, enabling AI systems to surpass traditional statistical methods.
