arXiv:2606. 25008v1 Announce Type: new Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute.
By Yizhou Liu, Jeff Gore
arXiv:2602. 03685v2 Announce Type: replace-cross Abstract: Training large language models (LLMs) is computationally expensive, partly because the loss exhibits slow power-law convergence whose origin remains debatable.
By Yizhou Liu, Ziming Liu, Cengiz Pehlevan, Jeff Gore
arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.
By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.
By Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima
arXiv:2606. 06238v1 Announce Type: new Abstract: We propose a statistical-field framework for text generated by large language models (LLMs), treating token embeddings as continuous spin variables on a one-dimensional chain.
By Huajian Ruan, Jinyang Li, Xingyu Guo, Lingxiao Wang
arXiv:2606. 29158v1 Announce Type: cross Abstract: Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs.
By Zaiwen Yang, Huaqing Zhang, Jing Xu, Jingzhao Zhang
The paper investigates neural scaling laws for RydbergGPT, an autoregressive transformer trained on qubit measurement data from Rydberg atom arrays. Near a critical point in the quantum system, the transformer’s loss scales with dataset size following a power‑law with a loss‑floor correction, whereas this relationship weakens away from criticality. By comparing entropy‑normalized mutual‑information two‑point functions of Rydberg data and natural‑language corpora, the authors find that near‑critical statistics resemble natural language more closely, suggesting that multi‑scale dependence underlies stable neural scaling and that scaling behaviour depends on the model–data pair.
By David S. Berman, Ying-Jer Kao, Roger G. Melko, Alexander G. Stapleton
The paper investigates how learning rate and batch size scale when pretraining dense large language models on English‑prevalent corpora, examining both jointly optimal and marginal evolutions across model capacity and data size. It explores the benefits of a Warmup‑Stable‑Decay learning‑rate schedule, assessing whether optimal hyperparameters transfer between stable and decay phases, and evaluates loss scaling forms that capture interactions between model capacity and dataset size. The study provides a baseline scaling procedure and releases the full set of pretraining runs for future OpenEuroLLM development.
By Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
arXiv:2606. 29139v1 Announce Type: new Abstract: We study how the next-token prediction of an autoregressive Transformer language model changes under small perturbations of earlier input token embeddings.
By Matthias Br\"andel, Stephan K\"ohler, Oliver Rheinbach