arXiv AI By Kevin Zhou, Lisa Alazraki, Kris Cao, Marek Rei

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Read the original on arXiv AI →

arXiv:2606. 07597v1 Announce Type: cross Abstract: Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Aug 26

Data Mixing as Mixture Experiment: Response Surface Methodology and Optimal Design for Large Language Model Pretraining

The article proposes treating large language model (LLM) data mixing as a classical mixture experiment, where data domains are components, token shares are proportions, and proxy-training runs serve as design points. Using sparse second‑order Scheffé response‑surface models, the authors construct model‑robust Σ‑optimal designs that efficiently identify optimal data mixtures and reveal strong interaction effects, especially between weak domains and web‑derived text. Empirical results on RegMix show that these designs recover mixture rankings while reducing proxy runs by about 25%, demonstrating that data mixing can be optimized through experimental design rather than solely prediction.

By Yicheng Mao, Hongru Du
arXiv AI
Aug 17

Scaling Domain Data Repetition in LLM Pretraining

arXiv:2608. 14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)).

By Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu, Shuran Zheng, Jingzhao Zhang
arXiv AI
Sep 7

Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases

The paper titled "Less Data, Faster Training: repeating smaller datasets speeds up learning via sampling biases" examines how training on a smaller dataset with repeated samples can reduce computational cost compared to using a larger dataset. Across various tasks, architectures, and optimizers, this phenomenon cannot be explained by existing theory. The authors attribute the speedup to layer‑wise growth driven by sampling biases, which is more pronounced with smaller datasets, and provide both theoretical analysis and empirical evidence to support this claim.

By Jingwen Liu, Ezra Edelman, Surbhi Goel, Bingbin Liu
arXiv AI
Jun 16

FastMix: Fast Data Mixture Optimization via Gradient Descent

arXiv:2606. 14971v1 Announce Type: cross Abstract: While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and post-training remains a significant open problem.

By Haoru Tan, Sitong Wu, Yanfeng Chen, Jun Xia, Ruobing Xie, Bin Xia, Xingwu Sun, Xiaojuan Qi