arXiv Machine Learning By Arda Fazla, Antesh Upadhyay, Ege C. Kaya, M. Berk Sahin, Abolfazl Hashemi

Why Does Adaptive Batching Help LLM Pretraining? A Perspective from Unbounded Variance

Read the original on arXiv Machine Learning →

The paper investigates why increasing batch size during large language model pretraining helps, challenging the common assumption of bounded stochastic gradient variance. By introducing a generalized Blum–Gladyshev noise model that allows variance to grow with a tunable exponent, the authors derive theoretical bounds on oracle complexity and design an adaptive batch scheduler that adjusts batch size dynamically. Experiments on OLMo2 models up to 1B parameters show that this scheduler achieves lower validation loss than fixed small or large batch training while using fewer iterations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Sep 24

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.

By Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou
arXiv AI
Jul 29

Bridging Compute- and Data-Optimal Pretraining

arXiv:2607. 25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data.

By Tian Qin, Kimia Hamidieh, David Alvarez-Melis