arXiv:2609.39595v1 Announce Type: new
Abstract: Practical Muon maintains momentum and performs a small, fixed number of Newton--Schulz iterations separately for each parameter matrix, often with a Ne...
By Hanyng Peng, Hui Wang, Yue Yu
The paper introduces Muon, an optimizer that uses a finite number of Newton‑Schulz iterations to approximate the polar factor for matrix‑valued parameters in large language model pretraining. It demonstrates that this finite iteration smooths the discontinuous polar map into a Lipschitz function of singular values, enabling a conversion from online learning regret to a stationarity guarantee in nonsmooth nonconvex optimization. The authors prove that a logarithmic depth in Newton‑Schulz suffices for convergence to stationary points, matching best‑known sample complexity bounds and extending the result to other spectral maps with similar smoothing properties.
By Mingyi Li, Taira Tsuchiya
arXiv:2608. 04607v1 Announce Type: cross Abstract: Stochastic gradient descent (SGD) optimization methods are the standard instruments for the training of deep neural networks (DNNs).
By Thang Do, Steffen Dereich, Arnulf Jentzen
arXiv:2606. 30509v1 Announce Type: new Abstract: Matrix factorization (i.
By Mark Rhee, Jamie Simon, Dhruva Karkada
arXiv:2509. 14562v4 Announce Type: replace Abstract: Large models recently are widely applied in machine learning, so efficient training of large models has received widespread attention.
By Feihu Huang, Yuning Luo, Songcan Chen
arXiv:2607. 20512v1 Announce Type: cross Abstract: The Muon optimizer reaches the grokking threshold on modular arithmetic faster than AdamW.
By Yufeng Wang
arXiv:2609.36692v1 Announce Type: cross
Abstract: Matrix optimizers have emerged as a promising direction, with Muon standing out as a prominent design. Revisiting Muon through its full-Gram represen...
By Zixuan Gong, Zeyu Gan, Jiaye Teng, Yong Liu
arXiv:2512. 04632v2 Announce Type: replace Abstract: Orthogonality-based optimizers, such as Muon, have recently shown strong performance across large-scale training and community-driven efficiency challenges.
By Thibaut Boissin (IRIT-MISFIT), Thomas Massena (DTIPG - SNCF, IRIT-MISFIT), Franck Mamalet (IRIT-MISFIT), Mathieu Serrurier (IRIT-MISFIT)
arXiv:2609.13677v1 Announce Type: cross
Abstract: Modern real application problems involve matrix-valued parameters, yet conventional optimizers treat them as vectors, thereby motivating matrix-aware...
By Lexiao Lai, Tianyi Lin, Jiayu Zhang
arXiv:2606. 21828v2 Announce Type: replace-cross Abstract: Neural operators are increasingly used to warm-start Newton solvers for nonlinear PDEs, on the premise that a low test error places the initial guess inside the basin of attraction.
By Jaemin Oh, Youngkyu Lee, Jerome Darbon, George Em Karniadakis
The limited-memory BFGS (L-BFGS) algorithm is a cornerstone of large-scale optimization due to its linear memory and computational costs. However, in ill-conditioned or non-convex landscapes, the implicit inverse Hessian approximation can suffer from an exploding condition number, leading to numerical instability and degraded convergence.
arXiv:2609.07597v1 Announce Type: cross
Abstract: Muon can be interpreted as optimizing a linear local objective over a spectral-norm ball. This gives a matrix-sign update that preserves the singular...
By Qiaozhe Zhang, Jun Sun, Yingzhuang Liu