Stabilizing Native Low-Rank LLM Pretraining
arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.
The paper introduces MSign, an optimizer designed to prevent training instability in large language models by restoring the stable rank of weight matrices. It identifies two precursors to gradient explosions—rapid stable rank decline and increased Jacobian alignment—and proves that these jointly cause exponential gradient growth. Experiments on models ranging from 5 M to 3 B parameters show that MSign stops training failures while adding less than 7.0% computational overhead.
arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.
arXiv:2606. 00888v1 Announce Type: cross Abstract: Dynamic Sparse Training (DST) offers a promising paradigm for improving the training and inference efficiency of deep neural networks; however, we find that in large language model training, DST can suffer from optimization instability, manifested as loss spikes after topology updates.
The paper investigates the often-overlooked scale vectors in large language models, showing that despite their tiny size they are crucial for pre‑training performance. The authors provide theoretical insights that scale vectors mainly aid optimization rather than expressivity, and they analyze how weight decay affects different normalization layers. Building on these findings, they propose lightweight improvements—branch‑specific heterogeneity, better placement, and magnitude‑direction reparameterization—that consistently reduce loss across a range of model sizes and training settings.
arXiv:2607. 09204v1 Announce Type: cross Abstract: Pretrained language models often exhibit structured weight spectra, suggesting that training may repeatedly produce similar layerwise and component-wise organization.
arXiv:2602.02848v2 Announce Type: replace Abstract: Advances in large language models have driven strong performance across many tasks, but their memory and compute costs still hinder deployment. SVD...
arXiv:2502. 11034v3 Announce Type: replace Abstract: Loss spikes remain a persistent obstacle in large-scale language model pretraining.
The paper investigates why the orthogonal optimiser Muon outperforms Adam in large language model pretraining by analysing the spectral properties of Transformer loss landscapes. It finds that Muon’s momentum buffers exhibit an anisotropic spectral profile with a volatile head and a tolerant bulk, enabling larger effective step sizes. Building on this insight, the authors propose Spectral‑Aware Muon (SAMuon) and a lightweight variant, which adjust the bulk scaling while keeping the head unchanged, achieving 13–24 % fewer training tokens than Muon without extra FLOPs.
The paper introduces a scalable Kronecker-based approximation that captures cross-layer interactions without storing the full Fisher matrix, making Hessian analysis feasible for billion-parameter language models. It identifies consistent vulnerability patterns, notably that value projection layers are the most sensitive and exhibit strong cross-layer correlations across various model families. Experiments on quantization, sparsification, inter-layer corruption, and fine-tuning show that the approximation correlates strongly with performance degradation and recovery, providing a practical tool for identifying fragile components and guiding compression and optimization strategies.
arXiv:2605. 09825v4 Announce Type: replace-cross Abstract: Why does full-pipeline FP4 training of large language models often diverge, even when forward activations and activation gradients remain stable?
arXiv:2607. 10803v1 Announce Type: cross Abstract: Understanding which parameters are influential in Large Language Models (LLMs) is central to improving their efficiency, reliability, and interpretability.
The paper addresses training instabilities in large language model pretraining, specifically output logit divergence that occurs near the end of training. By analyzing the geometry of output embeddings, the authors identify anisotropic embeddings as the root cause and propose Output Embedding Centering (OEC) as a mitigation strategy. OEC can be applied deterministically as μ‑centering or as a regularization loss μ‑loss, and experiments show both variants outperform the existing z‑loss method while matching logit soft‑capping in stability, even without weight tying. Additionally, μ‑loss is less sensitive to hyperparameter tuning than z‑loss.
arXiv:2606. 05516v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables memory-efficient fine-tuning of large language models (LLMs) using only forward passes, but it remains unclear how useful adaptation is distributed across layers.