arXiv AI

SPA: A Simple but Tough-to-Beat Baseline for Knowledge Injection

arXiv:2603. 22213v2 Announce Type: replace-cross Abstract: While large language models (LLMs) are pretrained on massive amounts of data, their knowledge coverage remains incomplete in specialized, data-scarce domains, motivating extensive efforts to study synthetic data generation for knowledge injection.

arXiv Machine Learning
Jul 31

How Can We Synthesize High-Quality Pretraining Data? A Systematic Study of Prompt Design, Generator Model, and Source Data

arXiv:2604. 13977v2 Announce Type: replace-cross Abstract: Synthetic data is a standard component in training large language models, yet systematic comparisons across design dimensions, including rephrasing strategy, generator model, and source data, remain absent.

By Joel Niklaus, Atsuki Yamaguchi, Michal \v{S}tef\'anik, Guilherme Penedo, Hynek Kydl\'i\v{c}ek, Elie Bakouch, Lewis Tunstall, Edward Emanuel Beeching, Thibaud Frere, Colin Raffel, Leandro von Werra, Thomas Wolf
arXiv AI
Sep 10

Generating Pretraining Tokens from Organic Data for Data-Bound Scaling

The paper introduces SynPro, a synthetic data generation framework that augments limited organic text for large language model pretraining by applying rephrasing and reformatting operations. SynPro’s generators are optimized with reinforcement learning rewards for quality, faithfulness, and data influence, and are updated continuously as training plateaus. Experiments on 400M, 1.1B, and 2B models show that SynPro can unlock 3.4–5.2× the effective tokens of standard repetition, even outperforming a non‑data‑bound oracle at larger scales.

By Zichun Yu, Chenyan Xiong
arXiv AI
Jun 16

Data Augmentations for Data-Constrained Language Model Pretraining

arXiv:2606. 16246v1 Announce Type: cross Abstract: As AI labs approach a data ceiling where compute capacity outpaces the rate of new high-quality text generation, language model pretraining is shifting toward a data-constrained, compute-abundant regime that demands productive multi-epoch training on fixed corpora.

By Michael K. Chen, Xikun Zhang, Zhen Wang
arXiv Machine Learning
4d ago

It's All Training: A Fully Synthetic Single-Stage Recipe for LLMs

arXiv:2609.37891v1 Announce Type: cross Abstract: Current pre-training datasets are derived from web crawls, with all their issues, and were not designed to support mid- and post-training pipelines--...

By Pierre-Carl Langlais, Pieter Delobelle, Yannick Detrois, Pavel Chizhov, Carlos Rosas-Hinostroza, Neil Si Smail, Benjamin Burtin, Hanna Shcharbakova, Ivan Yamshchikov, Anastasia Stasenko
arXiv AI
Aug 20

Demystifying Training-Time Augmentation for Data-Constrained Language Model Pretraining

The paper investigates training-time data augmentation as a regularizer for autoregressive language model pretraining in data‑constrained, compute‑abundant settings. It introduces three orthogonal augmentation categories—token‑level noise, sequence permutations, and target offset prediction—and shows through systematic ablations that each category delays overfitting and reduces validation loss, with random token replacement performing best individually. Combining augmentation categories further lowers the minimum validation loss, demonstrating that such augmentations mitigate data inefficiency in autoregressive pretraining.

By Michael K. Chen, Xikun Zhang, Fan Bai, Zhengding Hu, Zhen Wang
arXiv Machine Learning
Sep 10

RePro: Training Language Models to Faithfully Recycle the Web for Pretraining

RePro is a web‑recycling technique that trains a small language model (as little as 1 B parameters) with reinforcement learning to produce high‑quality, faithful rephrasings of pretraining data. The method uses one quality reward and three faithfulness rewards to preserve core semantics and structure while converting organic data into better training examples. Experiments show that RePro boosts downstream accuracy by 3.7–14.5 % over organic‑only baselines and improves data efficiency 2–3×, outperforming prior prompting‑based recycling approaches.

By Zichun Yu, Chenyan Xiong
arXiv Machine Learning
Aug 31

Learning a Size-Weight Frontier for Synthetic-Augmented Inference

The paper introduces a framework for synthetic‑augmented inference that balances the number of synthetic observations with their assigned weight. It defines a size‑weight frontier, estimating for each weight the maximum synthetic sample size that still guarantees target task‑marginal coverage for all smaller sizes. The authors provide finite‑sample coverage guarantees for configurations on or below this frontier and demonstrate that, when applied to augment opinion survey data with large language model responses, the method achieves the desired coverage while significantly tightening confidence intervals.

By Chengpiao Huang, Kaizheng Wang