arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.
By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
arXiv:2606. 25008v1 Announce Type: new Abstract: Neural scaling laws describe how pre-training loss decays as power laws with training time, model size, and compute.
By Yizhou Liu, Jeff Gore
arXiv:2606. 29158v1 Announce Type: cross Abstract: Learning-rate transfer can reduce the cost of training large language models: instead of sweeping learning rates at target scale, practitioners extrapolate from smaller runs.
By Zaiwen Yang, Huaqing Zhang, Jing Xu, Jingzhao Zhang
arXiv:2512. 22088v3 Announce Type: replace-cross Abstract: The scaling law, a cornerstone of Large Language Model (LLM) development, predicts improvements in model performance with increasing computational resources.
By Chiwun Yang
The paper introduces Power‑Law Entropy Search (PLES), a computational‑cost‑aware acquisition function that uses multi‑fidelity Bayesian optimization to efficiently estimate optimal hyperparameter scaling laws for large language model training. PLES focuses on reducing the overall uncertainty of scaling law estimates rather than optimizing a single objective, selecting configurations that maximize uncertainty reduction per unit computational cost. Experiments on synthetic benchmarks, surrogate models, and real LLM pre‑training runs show that PLES converges to accurate scaling laws using less than one‑tenth of the computational budget required by conventional grid search and other baselines.
By Zhiliang Chen, Sebastian Ament, David Eriksson, Maximilian Balandat, Eytan Bakshy, Jihao Andreas Lin
The paper investigates how learning rate and batch size scale when pretraining dense large language models on English‑prevalent corpora, examining both jointly optimal and marginal evolutions across model capacity and data size. It explores the benefits of a Warmup‑Stable‑Decay learning‑rate schedule, assessing whether optimal hyperparameters transfer between stable and decay phases, and evaluates loss scaling forms that capture interactions between model capacity and dataset size. The study provides a baseline scaling procedure and releases the full set of pretraining runs for future OpenEuroLLM development.
By Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
arXiv:2602. 02712v2 Announce Type: replace Abstract: A popular approach to post-training control of large language models (LLMs) is the steering of intermediate latent representations.
By Magamed Taimeskhanov, Samuel Vaiter, Damien Garreau
arXiv:2605. 13026v2 Announce Type: replace-cross Abstract: Masked diffusion models (MDMs) have emerged as a promising alternative to autoregressive models (ARMs) for language modeling.
By Chunsan Hong, Sanghyun Lee, Chieh-Hsin Lai, Satoshi Hayakawa, Yuhta Takida, Yuki Mitsufuji, Seungryong Kim, Jong Chul Ye
arXiv:2512.00763v2 Announce Type: replace-cross
Abstract: Adaptive and non-Euclidean optimizers often outperform Euclidean methods such as stochastic gradient descent (SGD) in language modeling by a...
By Robin Yadav, Shuo Xie, Tianhao Wang, Zhiyuan Li
The paper proposes an information-weighted cross‑entropy loss that rescales token contributions using TF‑IDF statistics, thereby emphasizing semantically informative tokens and down‑weighting ubiquitous ones. Experiments on five decoder‑only language models (1.1B–13B parameters) show consistent reductions in memorized substring length while maintaining perplexity and downstream performance. The method is architecture‑agnostic, adds less than 3% computational overhead, and can be integrated into existing training pipelines.
By Zhijian Li, Stefan Larson, Kevin Leach
arXiv:2607. 14306v1 Announce Type: new Abstract: In this paper, we study the connection between an LLM's output distribution and the data used to train it.
By Zachary Izzo