Optimizer-dependent training dynamics converge to the same one-third optimal data scaling
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.27581v1 Announce Type: new Abstract: Step Law gives power-law formulas for the optimal peak learning rate eta* and batch size B* when pre-training language models. It was calibrated on mod...
arXiv:2609.37535v1 Announce Type: new Abstract: In the softmax output layer, a rare token receives a small positive logit gradient on most steps and a much larger negative gradient on the few steps w...
The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.
arXiv:2602.06797v3 Announce Type: replace-cross Abstract: We study optimal learning rate (LR) schedules under the functional scaling law (FSL) framework (Li et al., 2025), which decomposes training d...
arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.
arXiv:2609. 25710v1 Announce Type: cross Abstract: The statistical accuracy of neural networks depends on both their approximation power and the complexity of the class fitted from data.