arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.
By Ali Parviz, Gal Mishne, Alex Cloninger
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.
arXiv:2606. 08388v1 Announce Type: new Abstract: Muon replaces a matrix gradient $G=U\Sigma V^\top$ by its polar factor $UV^\top$.
By Pierfrancesco Beneventano, Mahmoud Abdelmoneum, Tomaso Poggio
AF‑Muon is an AdamW‑free extension of the Muon optimizer that retains Muon’s matrix update for hidden weights while applying a support‑aware finite‑cap linear minimization oracle to tied vocabulary tables and an RMS‑normalized update for one‑dimensional auxiliary parameters. This design eliminates second‑moment state, reducing optimizer‑state memory by about 20% compared to Hybrid Muon. Across nine tied‑token settings—including decoder‑only language models, T5‑style encoder‑decoders, and ImageGPT‑style variants—AF‑Muon consistently improves mean validation loss and perplexity over both Hybrid Muon and a SCION‑style Sign endpoint, with robust gains confirmed by long‑horizon runs and hyperparameter studies.
By Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu
arXiv:2507. 01598v5 Announce Type: replace Abstract: Muon, a recently proposed optimizer that leverages the inherent matrix structure of neural network parameters, has demonstrated strong empirical performance, indicating its potential as a successor to standard optimizers such as AdamW.
By Naoki Sato, Hiroki Naganuma, Hideaki Iiduka
arXiv:2607. 19771v1 Announce Type: cross Abstract: Muon and related matrix-sign optimizers are increasingly used to pre-train large language models, but their effect on the internal geometry of individual weight matrices is not well understood.
By Jiachun Li