arXiv Machine Learning

Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum

arXiv Machine Learning
Aug 28

Muon with Finite Newton-Schulz: The Smoothing Benefit in Nonsmooth Nonconvex Optimization

The paper introduces Muon, an optimizer that uses a finite number of Newton‑Schulz iterations to approximate the polar factor for matrix‑valued parameters in large language model pretraining. It demonstrates that this finite iteration smooths the discontinuous polar map into a Lipschitz function of singular values, enabling a conversion from online learning regret to a stationarity guarantee in nonsmooth nonconvex optimization. The authors prove that a logarithmic depth in Newton‑Schulz suffices for convergence to stationary points, matching best‑known sample complexity bounds and extending the result to other spectral maps with similar smoothing properties.

By Mingyi Li, Taira Tsuchiya
arXiv Machine Learning
Sep 11

Musec: MomentUm SpEctral Clipping for Stable Muon-type Training

Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.

By Zhuanghua Liu, Menglian Wang, Luo Luo
arXiv AI
6d ago

Convergence guarantees for Muon: New parameter regimes and generalizations

The paper presents the first asymptotic convergence guarantees for the Muon algorithm, showing that with suitable hyperparameters the iterates satisfy ≠≠ ∥∇f(x_k)∥ → 0 and, under a global Polyak-ℒojasiewicz condition, the function values converge linearly. It reveals that Muon’s implicit regularization acts as a bounded preconditioner, framing Muon as a preconditioned Polyak heavy‑ball method and enabling a Lyapunov analysis. Building on this insight, the authors introduce Muesterov, a Nesterov‑based variant, and prove it shares the same convergence guarantees, extending the theory beyond the heavy‑ball setting; numerical experiments on a scalar cross‑entropy problem and preliminary nanoGPT simulations support the theoretical findings.

By Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag
arXiv Machine Learning
Jul 17

Muse: Representation Geometry of Muon Beyond Normalized Momentum

arXiv:2607. 14536v1 Announce Type: new Abstract: Muon-style optimizers apply a polar map to matrix momentum, but their updates also depend on the representation of each parameter block before orthogonalization.

By Da Chang, Qiankun Shi, Lvgang Zhang, Di He, Yaoshuai Ma, Ganzhao Yuan, Yongxiang Liu
arXiv Machine Learning
Sep 11

DP-Muon: Differentially Private Optimization via Matrix-Orthogonalized Momentum

The paper introduces DP-Muon, a differentially private optimization method that incorporates matrix‑orthogonalized momentum. It employs standard global per‑example clipping and releases a single Gaussian‑noised gradient per step, with matrix and auxiliary updates treated as post‑processing. The authors analyze the mean distortion introduced when fresh Gaussian noise passes through a nonlinear matrix map, deriving exact Gaussian heat identities and showing that for a smooth Newton‑Schulz map, the conditional output bias is reduced from second to fourth order in the noise scale. They also establish matrix‑block stationarity bounds, quantify orthogonalization error, and provide criteria for improving the upper bound, while a separate inequality captures the impact of auxiliary Adam updates. Experiments on GPT‑2 at various privacy targets demonstrate that DP‑Muon configurations outperform Adam baselines in test negative log‑likelihood.

By Jihwan Kim, Chenglin Fan
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar