arXiv:2406. 05335v3 Announce Type: replace-cross Abstract: Generation of text and speech in natural languages can be modeled as a stochastic process.
By Kai Nakaishi, Yoshihiko Nishikawa, Koji Hukushima
arXiv:2609.36455v1 Announce Type: new
Abstract: Large language models are thought to represent features by vectors in a hidden space of dimension given by the model's width. Superposition, in which m...
By Lihao Guo, Yizhou Liu, Jeff Gore
The paper investigates how learning rate and batch size scale when pretraining dense large language models on English‑prevalent corpora, examining both jointly optimal and marginal evolutions across model capacity and data size. It explores the benefits of a Warmup‑Stable‑Decay learning‑rate schedule, assessing whether optimal hyperparameters transfer between stable and decay phases, and evaluates loss scaling forms that capture interactions between model capacity and dataset size. The study provides a baseline scaling procedure and releases the full set of pretraining runs for future OpenEuroLLM development.
By Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
arXiv:2602. 07488v3 Announce Type: replace-cross Abstract: Despite the fact that experimental neural scaling laws have substantially guided empirical progress in large-scale machine learning, no existing theory can quantitatively predict the exponents of these important laws for any modern LLM trained on any natural language dataset.
By Francesco Cagnetta, Allan Ravent\'os, Surya Ganguli, Matthieu Wyart
arXiv:2506.17871v4 Announce Type: replace-cross
Abstract: Despite their impressive capabilities, aligned large language models (LLMs) often generate outputs that lack diversity. What drives this cons...
By Chenghao Yang, Sida Li, Ari Holtzman
arXiv:2606. 07559v1 Announce Type: cross Abstract: Fine-tuning a language model on contexts whose correct completion has a near-synonym competitor often fails silently.
By Vaibhav Prakash, Jayasri Dontabhaktuni