arXiv AI

How Does Local Landscape Geometry Evolve in Language Model Pre-Training?

The paper investigates how the local landscape geometry of language model pre‑training evolves, identifying two distinct phases. In Phase I, the landscape starts sharp, causing instability and loss plateaus at high learning rates, which explains the need for learning‑rate warmup and suggests longer warmups for larger peak rates. In Phase II, the geometry is governed by gradient noise scale, revealing a depth‑flatness trade‑off that motivates a dynamic batch‑size scheduler that starts small and grows later in training.

arXiv Machine Learning
Aug 31

Deriving Scaling Laws for OpenEuroLLM Models: Learning Rate, Batch Size and Loss

The paper investigates how learning rate and batch size scale when pretraining dense large language models on English‑prevalent corpora, examining both jointly optimal and marginal evolutions across model capacity and data size. It explores the benefits of a Warmup‑Stable‑Decay learning‑rate schedule, assessing whether optimal hyperparameters transfer between stable and decay phases, and evaluates loss scaling forms that capture interactions between model capacity and dataset size. The study provides a baseline scaling procedure and releases the full set of pretraining runs for future OpenEuroLLM development.

By Niccol\`o Ajroldi, Diana Alexandra Onutu, Haider Al-Tahan, J\"org Franke, Sampo Pyysalo, Jenia Jitsev, Aaron Klein
arXiv Machine Learning
5d ago

Towards Understanding Momentum Acceleration in River-Valley Loss Landscape

The paper investigates how momentum influences optimization in a river‑valley loss landscape, where a low‑loss manifold is surrounded by steep orthogonal directions. It shows that momentum stabilizes large learning rates that vanilla gradient descent cannot tolerate, enabling faster progress along the river. The study also finds that in very flat, slowly spinning rivers, momentum itself does not directly accelerate tracking, but the larger permissible learning rate does.

By Miao Lu, Zeyu Bian, Kaiyue Wen, Beining Wu, Siyu Chen, Tianhao Wang, Zhiyuan Li
arXiv Machine Learning
Jun 30

Why Do We Need Warm-up? A Theoretical Perspective

arXiv:2510. 03164v2 Announce Type: replace Abstract: Learning rate warm-up -- increasing the learning rate at the beginning of training -- has become a ubiquitous heuristic in modern deep learning, yet its theoretical foundations remain poorly understood.

By Foivos Alimisis, Rustem Islamov, Aurelien Lucchi
arXiv Machine Learning
Jul 15

AMUSE: Anytime Muon with Stable Gradient Evaluation

arXiv:2605. 22432v2 Announce Type: replace Abstract: Modern deep learning commonly relies on AdamW with prescribed learning rate schedules, but recent works challenge both components: Schedule-Free optimization removes explicit schedules via iterate averaging, and Muon improves the update geometry by orthogonalizing momentum for matrix parameters.

By Jueun Kim, Baekrok Shin, Jihun Yun, Beomhan Baek, Minhak Song, Chulhee Yun
arXiv AI
Aug 25

How Should LLMs Consume High-Quality Data? Optimal Data Scheduling via Quality-Aware Functional Scaling Laws

The paper proposes a theoretical framework for scheduling high‑quality data in large language model training by extending functional scaling laws to account for time‑varying data quality. It identifies two regimes—noise‑limited and signal‑limited—where high‑quality data should be used differently, and introduces a Drop‑Stable‑Rampup training schedule that adjusts batch size at the quality transition. Experiments on 15B MoE and 600M dense models show significant accuracy gains over conventional decay schedules across multiple benchmarks.

By Zhitao Zhu, Xili Wang, Shizhe Wu, Jiawei Fu, Xiaoqing Liu