arXiv Machine Learning By Vincent Chen, Starrick Liu, Regis Cheng, Dance Yang, Shalfun Li, Ryan Yu, Lucy Liang, Hang Su, Roy Gan, Hao Wang, Qian Wang

DMuon: Efficient Distributed Muon Training with Near-Adam Overhead

Read the original on arXiv Machine Learning →

arXiv:2606. 27153v1 Announce Type: cross Abstract: Matrix-orthogonalization-based optimizers, exemplified by Muon, have demonstrated strong convergence behavior across a wide range of modern deep learning workloads.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv Machine Learning
Sep 23

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.

By Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai
Hugging Face Trending Papers
Jul 14

Reassessing Muon for Matrix Factorization

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.

arXiv Machine Learning
1d ago

AF-Muon: An AdamW-Free Muon Optimizer for Tied-Embedding Models

AF‑Muon is an AdamW‑free extension of the Muon optimizer that retains Muon’s matrix update for hidden weights while applying a support‑aware finite‑cap linear minimization oracle to tied vocabulary tables and an RMS‑normalized update for one‑dimensional auxiliary parameters. This design eliminates second‑moment state, reducing optimizer‑state memory by about 20% compared to Hybrid Muon. Across nine tied‑token settings—including decoder‑only language models, T5‑style encoder‑decoders, and ImageGPT‑style variants—AF‑Muon consistently improves mean validation loss and perplexity over both Hybrid Muon and a SCION‑style Sign endpoint, with robust gains confirmed by long‑horizon runs and hyperparameter studies.

By Arash Lagzian, Paniz Halvachi, Junming Zhang, Zhouhan Lin, Dianbo Liu
arXiv AI
Jul 24

SOAP, Muon, and Beyond: Pushing LLM Pretraining Scales

arXiv:2607. 20548v1 Announce Type: cross Abstract: Higher-order optimizers such as Muon and SOAP offer faster convergence than AdamW, but their computational cost and numerical stability challenges have limited adoption at scale.

By Mikail Khona, Aditya Vavre, Boxiang Wang, Deyu Fu, Hao Wu, Mike Chrzanowski, Bryan Catanzaro, Dheevatsa Mudigere, Jeff Pool, Michael Lightstone, Mohammad Shoeybi, Mostofa Patwary, Nima Tajbakhsh, Tijmen Blankevoort