arXiv:2510. 14717v2 Announce Type: replace-cross Abstract: Increasing the batch size during training -- a ''batch ramp'' -- is a promising strategy to accelerate large language model pretraining.
By Alexandru Meterez, Depen Morwani, Jingfeng Wu, Costin-Andrei Oncescu, Cengiz Pehlevan, Sham Kakade
arXiv:2610. 00436v1 Announce Type: new Abstract: Online batch selection fine-tunes a language model on the most useful part of each candidate batch.
By Hongyu Chen, Xinyi Luo, Ming Zhao, Lin Tang, Zihan Xu, Jing Li, Yuxuan Wang, Haoran Deng, Wei Zhang
The paper presents empirical scaling laws for autoregressive language models, linking prediction loss to model size, data size, and compute, and investigates their theoretical basis using a teacher–student linear RNN framework. In this tractable setting, a stable latent linear RNN generates trajectories while a sketched linear recurrent student is trained via full‑batch WSD gradient descent on next‑token prediction. The study derives explicit approximation, optimization, and statistical scaling laws that depend on the sketch dimension, number of trajectories, and trajectory length, revealing how different power‑law exponents for innovation and initialization covariances affect the rates and crossovers between regimes.
By Ziyan Chen, Zhongzhu Zhou, Peilin Liu, Ding-Xuan Zhou
arXiv:2606. 25086v1 Announce Type: new Abstract: Many modern Language Model (LM) pipelines return an averaged model, such as an exponential moving average of the training iterates, rather than the final iterate itself.
By Kwok Chun Au, Adam Block
arXiv:2607. 25271v1 Announce Type: cross Abstract: Classical compute-optimal scaling laws assume an unbounded supply of fresh pretraining data, yet pretraining is increasingly entering a regime in which compute grows faster than the availability of high-quality data.
By Tian Qin, Kimia Hamidieh, David Alvarez-Melis
arXiv:2606. 19179v1 Announce Type: cross Abstract: Stochastic momentum methods such as heavy ball (HB), Nesterov momentum, and variants of Accelerated SGD (ASGD) [Kidambi et al.
By Depen Morwani, Alexandru Meterez, Pranav Nair, Sham Kakade
The paper proposes a theoretical framework for scheduling high‑quality data in large language model training by extending functional scaling laws to account for time‑varying data quality. It identifies two regimes—noise‑limited and signal‑limited—where high‑quality data should be used differently, and introduces a Drop‑Stable‑Rampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.
By Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu
The paper investigates how the local landscape geometry of language model pre‑training evolves, identifying two distinct phases. In Phase I, the landscape starts sharp, causing instability and loss plateaus at high learning rates, which explains the need for learning‑rate warmup and suggests longer warmups for larger peak rates. In Phase II, the geometry is governed by gradient noise scale, revealing a depth‑flatness trade‑off that motivates a dynamic batch‑size scheduler that starts small and grows later in training.
By Zhanpeng Zhou, Yuhan Sun, Bingrui Li, Jinbo Wang, Huaijin Wu, Lei Wu, Junchi Yan
arXiv:2608. 11361v1 Announce Type: new Abstract: Tokenizer vocabulary size is a foundational design choice in large language model (LLM) infrastructure, yet it is typically fixed at training time based on convention rather than deployment analysis.
By Rima Mittal, Ankit Gubrani, Satyanarayana Kakollu
arXiv:2608.29296v1 Announce Type: cross
Abstract: Larger batches reduce the variance of stochastic gradients per update and are therefore often expected to accelerate training. Yet whether this stati...
By Ziniu Li, Jinbo Wang, Guanhua Huang, Feiyuan Zhang, Pengbo Li, Alex Chen
arXiv:2602. 03001v2 Announce Type: replace-cross Abstract: To maximize hardware utilization, modern machine learning systems typically employ large constant or manually tuned batch size schedules, relying on heuristics that are brittle and costly to tune.
By Hiroki Naganuma, Shagun Gupta, Youssef Briki, Ioannis Mitliagkas, Irina Rish, Parameswaran Raman, Hao-Jun Michael Shi
arXiv:2609.30499v1 Announce Type: new
Abstract: Uniform noise-moment bounds exclude stochastic gradients whose variability increases with the iterate. We study ordinary, single-sample stochastic grad...
By Wei Biao Wu