arXiv AI By Rui Dai, Shuran Zheng

Explaining Data Mixing Scaling Laws

Read the original on arXiv AI →

arXiv:2606. 08167v1 Announce Type: cross Abstract: Recent research has established empirical scaling laws to predict model performance on multi-domain data mixtures.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

arXiv AI
2d ago

Scaling Domain Data Repetition in LLM Pretraining

arXiv:2608. 14071v1 Announce Type: new Abstract: As large language models scale, their training-token budgets must also increase to maintain an appropriate tokens-per-parameter ratio (\(\mathrm{TPP}\)).

By Jingwei Li, Xinran Gu, Rui Dai, Xintong Hao, Chengyin Xu, Yan Wu, Shuran Zheng, Jingzhao Zhang