arXiv Machine Learning

From Spectra to Joint Schedules in LLM Pre-training: 3+3(+2) Scaling-Law Regimes

Hugging Face Trending Papers
Sep 2

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

The paper investigates how momentum methods affect large‑batch training in a one‑pass setting using power‑law kernel regression. It derives critical learning rates for SGD, Polyak, and Nesterov, and shows how these rates depend on batch size, momentum, and model capacity. The authors provide scaling laws for risk dynamics, a three‑regime batch‑size phase diagram, and demonstrate that Polyak increases the critical batch size while Nesterov improves data efficiency in the large‑batch regime.

arXiv Machine Learning
Sep 3

Momentum in large-batch training: Polyak enlarges the critical batch size, Nesterov improves data efficiency

The paper investigates how momentum methods affect large‑batch training in a one‑pass setting using power‑law kernel regression. It derives critical learning rates for SGD, Polyak, and Nesterov, and shows how these rates depend on batch size, momentum, and a capacity exponent. The authors then analyze risk dynamics, optimize final‑step risk under a fixed data budget, and present a three‑regime batch‑size phase diagram that highlights Polyak’s ability to enlarge the critical batch size and Nesterov’s superior data efficiency in the large‑batch regime.

By Jia-Nan Wang, Zixun Huang, Kairui Li, Lei Wu
arXiv AI
Sep 15

One Spectrum, Two Resources: Data-Memory Scaling in Autoregressive Prediction

The paper investigates how much learned memory is required to leverage additional data in autoregressive prediction models. It introduces a predictive‑energy spectrum that jointly governs data and memory scaling, proving a minimax law that links the number of prediction blocks and the size of the learned state to this spectrum. The authors demonstrate that optimal bit allocation and masked query‑key attention mechanisms realize this law, and they provide experimental evidence across multiple pretrained‑model scales.

By Chiwun Yang, Xiaoyu Li
arXiv Machine Learning
Sep 24

Linear RNN Scaling Laws: When Longer Sequences Beat More Sequences

The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.

By Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou
arXiv Machine Learning
Sep 23

A Spectral Theory of Grokking: Weight Decay induces Feature Learning

The paper presents a spectral theory explaining the phenomenon of grokking, where an initial fit to training data is followed by a delayed improvement in generalization. It shows that for homogeneous networks trained with squared loss and L₂ weight decay, residuals after memorization influence the neural tangent kernel (NTK) dynamics, leading to a transition from lazy to rich learning. The theory predicts that grokking timescales depend on the product of learning rate and weight decay, and that stronger decay can halt fitting, with empirical validation on modular addition tasks using MLPs and Transformers.

By Lenz Pracher, Pascal de Jong, Oskar Lieshaus, Alan Jeffares, Steffen Rulands