arXiv AI By Kevin Zhou, Lisa Alazraki, Kris Cao, Marek Rei

Repetition Mismatch: Why Data Mixture Experiments Don't Scale and How to Fix Them

Read the original on arXiv AI →

arXiv:2606. 07597v1 Announce Type: cross Abstract: Pre-training data mixtures are commonly tuned by running small-scale experiments and extrapolating to the target training budget.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
3d ago

Scaling Domain Data Repetition in LLM Pretraining

arXiv:2608. 14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)).

By Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu, Shuran Zheng, Jingzhao Zhang
arXiv AI
Jun 16

FastMix: Fast Data Mixture Optimization via Gradient Descent

arXiv:2606. 14971v1 Announce Type: cross Abstract: While large and diverse datasets have driven recent advances in large models, identifying the optimal data mixture for pre-training and post-training remains a significant open problem.

By Haoru Tan, Sitong Wu, Yanfeng Chen, Jun Xia, Ruobing Xie, Bin Xia, Xingwu Sun, Xiaojuan Qi