arXiv Machine Learning

Inner Momentum for Differentially Private Muon

The paper introduces Inner Momentum (IM) for differentially private Muon training. It addresses distortion caused by per-example gradient clipping by averaging each example’s Muon gradient over the current and recent models before clipping, thereby bounding clipping-induced distortion. Experiments on private GPT‑2 fine‑tuning show that DP‑Muon‑IM consistently improves BLEU and ROUGE‑L scores and reduces polar error compared to standard DP‑Muon.

arXiv Machine Learning
Sep 11

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.

By Zhuanghua Liu, Menglian Wang, Luo Luo
arXiv Machine Learning
Sep 11

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

The paper introduces DP-Muon, a differentially private optimization method that incorporates matrix‑orthogonalized momentum. It employs standard global per‑example clipping and releases a single Gaussian‑noised gradient per step, with matrix and auxiliary updates treated as post‑processing. The authors analyze the mean distortion introduced when fresh Gaussian noise passes through a nonlinear matrix map, deriving exact Gaussian heat identities and showing that for a smooth Newton‑Schulz map, the conditional output bias is reduced from second to fourth order in the noise scale. They also establish matrix‑block stationarity bounds, quantify orthogonalization error, and provide criteria for improving the upper bound, while a separate inequality captures the impact of auxiliary Adam updates. Experiments on GPT‑2 at various privacy targets demonstrate that DP‑Muon configurations outperform Adam baselines in test negative log‑likelihood.

By Jihwan Kim, Chenglin Fan
arXiv Machine Learning
Jul 17

Muse: Representation Geometry of Muon Beyond Normalized Momentum

arXiv:2607. 14536v1 Announce Type: new Abstract: Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization.

By Da Chang, Qiankun Shi, Lvgang Zhang, Di He, Yaoshuai Ma, Ganzhao Yuan, Yongxiang Liu
arXiv AI
2d ago

Does Muon Need Fine-Grained Spectral Shaping?

The paper investigates how much spectral detail is necessary for the Muon optimizer, which traditionally uses a flat spectral profile. By analyzing singular modes, the authors find that most modes lie below a noise edge yet align positively with the gradient, leading them to propose BulkBoost—a two‑band spectral reweighting framework that adjusts bulk and spike gains while preserving matrix norms. Experiments across various pre‑training settings show that this low‑dimensional reweighting matches or surpasses fine‑grained spectral profiles, achieving notable loss reductions.

By Meher Chaitanya, Tianyi Zhou, Aristides Gionis
arXiv Machine Learning
Oct 2

The Row Normalization Puzzle in Muon

arXiv:2609.39114v1 Announce Type: new Abstract: This paper examines how row-wise renormalization affects Muon, focusing on the gap between NorMuon's worst-case guarantees and its practical performanc...

By Jiayu Zhang, Tianyi Lin
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar