Convergence of Practical Muon with Finite Newton-Schulz Iterations and Nesterov Momentum
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The paper introduces Muon, an optimizer that uses a finite number of Newton‑Schulz iterations to approximate the polar factor for matrix‑valued parameters in large language model pretraining. It demonstrates that this finite iteration smooths the discontinuous polar map into a Lipschitz function of singular values, enabling a conversion from online learning regret to a stationarity guarantee in nonsmooth nonconvex optimization. The authors prove that a logarithmic depth in Newton‑Schulz suffices for convergence to stationary points, matching best‑known sample complexity bounds and extending the result to other spectral maps with similar smoothing properties.
arXiv:2608. 04607v1 Announce Type: cross Abstract: Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs).
arXiv:2606. 08783v1 Announce Type: cross Abstract: Orthogonalized momentum updates, as used in Muon-style optimizers, have recently shown strong empirical stability in large-scale deep learning.
arXiv:2609.13677v1 Announce Type: cross Abstract: Modern real application problems involve matrix-valued parameters, yet conventional optimizers treat them as vectors, thereby motivating matrix-aware...
Musec introduces MomentUm SpEctral Clipping, an optimizer-level, architecture‑agnostic technique that replaces Muon’s spectral flattening with selective spectral clipping to stabilize training. By clipping singular values above a threshold while preserving the momentum’s spectral structure, Musec addresses loss spikes and unbounded weight growth without requiring architecture‑specific changes. Soft Musec, an efficient implementation using smooth spectral saturation via coupled Newton‑Schulz iterations, offers convergence guarantees in nonconvex nonsmooth stochastic optimization and empirically improves stability across diverse learning rates and model sizes.
The paper presents the first asymptotic convergence guarantees for the Muon algorithm, showing that with suitable hyperparameters the iterates satisfy ≠≠ ∥∇f(x_k)∥ → 0 and, under a global Polyak-ℒojasiewicz condition, the function values converge linearly. It reveals that Muon’s implicit regularization acts as a bounded preconditioner, framing Muon as a preconditioned Polyak heavy‑ball method and enabling a Lyapunov analysis. Building on this insight, the authors introduce Muesterov, a Nesterov‑based variant, and prove it shares the same convergence guarantees, extending the theory beyond the heavy‑ball setting; numerical experiments on a scalar cross‑entropy problem and preliminary nanoGPT simulations support the theoretical findings.