arXiv Machine Learning

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

arXiv:2607. 29674v1 Announce Type: cross Abstract: SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget.

arXiv Machine Learning
Aug 20

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

LionMuon is a new optimizer that alternates between Lion’s sign-based updates and Muon’s spectral matrix-sign updates on a fixed period P, sharing a single dual-EMA momentum buffer. This design keeps the memory footprint the same as Lion and half that of AdamW while reducing the average iteration cost compared to Muon. Experiments on 124M, 355M, and 720M models show LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW across datasets and architectures, achieving lower validation loss with less compute.

By Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horv\'ath, Martin Tak\'a\v{c}, Aleksandr Beznosikov
arXiv Machine Learning
Aug 27

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

StoSignSGD is a new sign‑based optimization algorithm that injects structural stochasticity into the sign operator, ensuring unbiased updates. It resolves the divergence issues of traditional SignSGD on non‑smooth objectives, achieving optimal convergence rates in convex settings and improved complexity bounds in non‑convex, non‑smooth problems. Empirical results show that StoSignSGD is stable and efficient across large language model training, outperforming AdamW and SignSGD in low‑precision regimes (FP8 and FP4) and delivering speedups and accuracy gains on models ranging from OLMo2‑370M to 7B LLMs.

By Dingzhi Yu, Rui Pan, Yuxing Liu, Difan Zou, Tong Zhang
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
arXiv AI
Aug 25

A Physical Response-and-Memory Model for Muon Optimization

The paper introduces a physical response-and-memory model for the Muon optimizer, explaining its semi‑orthogonalized momentum update as the maximally dissipative direction under an output‑side safety budget. It treats the weight matrix as a responsive medium with internal stress, showing that momentum corresponds to accumulated stress whose relaxation occurs over multiple timescales—fast and slow. Based on this, the authors propose the Bi‑Maxwell optimizer, which uses a two‑timescale memory kernel and achieves target loss in fewer steps on a public large‑language‑model benchmark.

By Yinze Hu, Hongjun Xiang, Xingao Gong, Hongyu Yu
arXiv AI
6d ago

Convergence guarantees for Muon: New parameter regimes and generalizations

The paper presents the first asymptotic convergence guarantees for the Muon algorithm, showing that with suitable hyperparameters the iterates satisfy ≠≠ ∥∇f(x_k)∥ → 0 and, under a global Polyak-ℒojasiewicz condition, the function values converge linearly. It reveals that Muon’s implicit regularization acts as a bounded preconditioner, framing Muon as a preconditioned Polyak heavy‑ball method and enabling a Lyapunov analysis. Building on this insight, the authors introduce Muesterov, a Nesterov‑based variant, and prove it shares the same convergence guarantees, extending the theory beyond the heavy‑ball setting; numerical experiments on a scalar cross‑entropy problem and preliminary nanoGPT simulations support the theoretical findings.

By Arthur C. B. de Oliveira, Dhruv D. Jatkar, Guilherme S. Vicinansa, Eduardo D. Sontag