arXiv Machine Learning By Tongle Wu, Huanyu Dong, Ying Sun, Ziye Ma

MALT: Lightweight Curvature-Aware Muon via Diagonal Preconditioning

Read the original on arXiv Machine Learning →

arXiv:2608. 05088v1 Announce Type: new Abstract: Muon has recently emerged as a promising alternative to AdamW for language model pretraining by orthogonalizing momentum matrices using Newton-Schulz iterations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 31

Curvature-Conditioned Multiscale Momentum with Sphere Constraints for LLM Pretraining

The paper introduces a curvature‑conditioned multiscale momentum algorithm with sphere constraints to accelerate large‑language‑model pretraining. By applying a slow‑decay component for noise reduction and a fast‑decay component for curvature adaptation only along flat directions, the method improves training dynamics without causing parameter inflation. Experiments demonstrate significant speed‑ups for Muon across various architectures and model sizes, and the authors provide theoretical justification for the observed acceleration.

By Shuchen Zhu, Yuxin Fang, Mingze Wang, Kun Yuan
arXiv Machine Learning
Sep 23

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.

By Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai
arXiv Machine Learning
Jun 16

CacheMuon: Using Temporal Preconditioning To Approximate Polar Factor

arXiv:2606. 16371v1 Announce Type: new Abstract: Muon is an optimizer that computes updates using the polar factor of the momentum matrix and has shown strong empirical performance across a range of training settings.

By Bishnu Dev (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Sushil Bohara (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Martin Tak\'a\v{c} (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE), Samuel Horv\'ath (Mohamed bin Zayed University of Artificial Intelligence, Abu Dhabi, UAE)