arXiv AI

Does Muon Need Fine-Grained Spectral Shaping?

The paper investigates how much spectral detail is necessary for the Muon optimizer, which traditionally uses a flat spectral profile. By analyzing singular modes, the authors find that most modes lie below a noise edge yet align positively with the gradient, leading them to propose BulkBoost—a two‑band spectral reweighting framework that adjusts bulk and spike gains while preserving matrix norms. Experiments across various pre‑training settings show that this low‑dimensional reweighting matches or surpasses fine‑grained spectral profiles, achieving notable loss reductions.

arXiv Machine Learning
Aug 27

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

The paper investigates why the orthogonal optimiser Muon outperforms Adam in large language model pretraining by analysing the spectral properties of Transformer loss landscapes. It finds that Muon’s momentum buffers exhibit an anisotropic spectral profile with a volatile head and a tolerant bulk, enabling larger effective step sizes. Building on this insight, the authors propose Spectral‑Aware Muon (SAMuon) and a lightweight variant, which adjust the bulk scaling while keeping the head unchanged, achieving 13–24 % fewer training tokens than Muon without extra FLOPs.

By Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
arXiv Machine Learning
1d ago

Spectrally Targeted Muon

The paper introduces Spectrally Targeted Muon, an optimizer that orthogonalizes only singular values of update matrices above or below a threshold τ, thereby interpolating between normalized SGD and the full Muon optimizer. It uses Newton‑Schulz iteration on a shifted Gram matrix to avoid SVD, and evaluates the method on CIFAR‑10 and NanoGPT, tracking effective rank and a new alignment metric. The study finds that orthogonalizing small singular values is crucial, that shrinking large singular values keeps the parameter spectrum flat, and that AdamW differs mainly in the speed of structure formation compared to Muon.

By Vishrut Goyal, Rohan Ramkumar
arXiv AI
Jun 12

LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold

arXiv:2606. 12921v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensitive to initialization choices, its optimal learning rates transfer poorly across ranks, and it often fails to beat dense baselines.

By Franz Louis Cesista, Katherine Crowson, C\'edric Simal, Stella Biderman
arXiv Machine Learning
Sep 11

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.

By Zhuanghua Liu, Menglian Wang, Luo Luo