The Spectral Dynamics and Noise Geometry of Muon
arXiv:2606. 08388v1 Announce Type: new Abstract: Muon replaces a matrix gradient $G=U\Sigma V^\top$ by its polar factor $UV^\top$.
The paper investigates how much spectral detail is necessary for the Muon optimizer, which traditionally uses a flat spectral profile. By analyzing singular modes, the authors find that most modes lie below a noise edge yet align positively with the gradient, leading them to propose BulkBoost—a two‑band spectral reweighting framework that adjusts bulk and spike gains while preserving matrix norms. Experiments across various pre‑training settings show that this low‑dimensional reweighting matches or surpasses fine‑grained spectral profiles, achieving notable loss reductions.
arXiv:2606. 08388v1 Announce Type: new Abstract: Muon replaces a matrix gradient $G=U\Sigma V^\top$ by its polar factor $UV^\top$.
The paper investigates why the orthogonal optimiser Muon outperforms Adam in large language model pretraining by analysing the spectral properties of Transformer loss landscapes. It finds that Muon’s momentum buffers exhibit an anisotropic spectral profile with a volatile head and a tolerant bulk, enabling larger effective step sizes. Building on this insight, the authors propose Spectral‑Aware Muon (SAMuon) and a lightweight variant, which adjust the bulk scaling while keeping the head unchanged, achieving 13–24 % fewer training tokens than Muon without extra FLOPs.
arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.
The paper introduces Spectrally Targeted Muon, an optimizer that orthogonalizes only singular values of update matrices above or below a threshold τ, thereby interpolating between normalized SGD and the full Muon optimizer. It uses Newton‑Schulz iteration on a shifted Gram matrix to avoid SVD, and evaluates the method on CIFAR‑10 and NanoGPT, tracking effective rank and a new alignment metric. The study finds that orthogonalizing small singular values is crucial, that shrinking large singular values keeps the parameter spectrum flat, and that AdamW differs mainly in the speed of structure formation compared to Muon.
arXiv:2608. 02991v1 Announce Type: new Abstract: Matrix spectral optimizers reshape weight-update spectra but usually delegate vector-valued biases to a separate optimizer.
arXiv:2605. 17109v3 Announce Type: replace-cross Abstract: In recent years, Muon has emerged as the dominant method for training large language models, and transformers more broadly.
arXiv:2606. 12921v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensitive to initialization choices, its optimal learning rates transfer poorly across ranks, and it often fails to beat dense baselines.
arXiv:2605.07815v2 Announce Type: replace Abstract: Muon fixes the \emph{direction} of every matrix-valued update at the polar factor of its momentum, while each layer's step \emph{magnitude} is addr...
Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.
arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.
arXiv:2607. 23711v1 Announce Type: new Abstract: LoRA fine-tuning can create intruder dimensions: new leading singular vectors of the updated weight matrix $W+BA$ that are nearly orthogonal to all pretrained singular vectors and that drive catastrophic forgetting.
arXiv:2607. 20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.