arXiv AI

An Isotropy-Preserving Spectral Cap for Muon: Theory and Three Case Studies

arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.

arXiv Machine Learning
Aug 27

Spectral Allocation: Why Muon Outperforms Adam, and How to Improve Muon

The paper investigates why the orthogonal optimiser Muon outperforms Adam in large language model pretraining by analysing the spectral properties of Transformer loss landscapes. It finds that Muon’s momentum buffers exhibit an anisotropic spectral profile with a volatile head and a tolerant bulk, enabling larger effective step sizes. Building on this insight, the authors propose Spectral‑Aware Muon (SAMuon) and a lightweight variant, which adjust the bulk scaling while keeping the head unchanged, achieving 13–24 % fewer training tokens than Muon without extra FLOPs.

By Xiaodong Wu, Wenyi Yu, Chao Zhang, Philip Woodland
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
arXiv AI
Aug 25

A Physical Response-and-Memory Model for Muon Optimization

The paper introduces a physical response-and-memory model for the Muon optimizer, explaining its semi‑orthogonalized momentum update as the maximally dissipative direction under an output‑side safety budget. It treats the weight matrix as a responsive medium with internal stress, showing that momentum corresponds to accumulated stress whose relaxation occurs over multiple timescales—fast and slow. Based on this, the authors propose the Bi‑Maxwell optimizer, which uses a two‑timescale memory kernel and achieves target loss in fewer steps on a public large‑language‑model benchmark.

By Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu
arXiv Machine Learning
Sep 11

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.

By Zhuanghua Liu, Menglian Wang, Luo Luo
arXiv Machine Learning
1d ago

AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models

AF‑Muon is an AdamW‑free extension of the Muon optimizer that retains Muon’s matrix update for hidden weights while applying a support‑aware finite‑cap linear minimization oracle to tied vocabulary tables and an RMS‑normalized update for one‑dimensional auxiliary parameters. This design eliminates second‑moment state, reducing optimizer‑state memory by about 20% compared to Hybrid Muon. Across nine tied‑token settings—including decoder‑only language models, T5‑style encoder‑decoders, and ImageGPT‑style variants—AF‑Muon consistently improves mean validation loss and perplexity over both Hybrid Muon and a SCION‑style Sign endpoint, with robust gains confirmed by long‑horizon runs and hyperparameter studies.

By Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu
arXiv Machine Learning
Sep 11

Toward a First-Principles Update Geometry for the Language-Model Head

The paper proposes a new update geometry for the language‑model head by treating the head and softmax as a single module and using Hilbert’s projective distance to measure functional change. It replaces the spectral norm with the Euclidean row diameter, derives a tractable RowNorm update rule, and demonstrates that RowNorm substantially reduces step diameters and Hilbert perturbations while only slightly increasing validation loss.

By Aditya Somasundaram, Charles Guille-Escuret, Alexander Moreno, Zhengzhong Liu, Eric Xing