arXiv Machine Learning By Maria Smirnova, Alexey Kravatskiy

Sign compression for Muon: SignMuon, MuonSign, and the Limits of Error Feedback

Read the original on arXiv Machine Learning →

arXiv:2607. 29674v1 Announce Type: cross Abstract: SignMuon compresses the Muon update to one bit per parameter by taking its elementwise sign, providing the most direct way to run a matrix-aware optimizer under an extremely low communication budget.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 20

LionMuon: Alternating Spectral and Sign Descent for Efficient Training

LionMuon is a new optimizer that alternates between Lion’s sign-based updates and Muon’s spectral matrix-sign updates on a fixed period P, sharing a single dual-EMA momentum buffer. This design keeps the memory footprint the same as Lion and half that of AdamW while reducing the average iteration cost compared to Muon. Experiments on 124M, 355M, and 720M models show LionMuon Pareto-dominates Muon, Lion, Signum, and AdamW across datasets and architectures, achieving lower validation loss with less compute.

By Arman Bolatov, Artem Riabinin, Nikita Kornilov, Andrey Veprikov, Samuel Horv\'ath, Martin Tak\'a\v{c}, Aleksandr Beznosikov
arXiv Machine Learning
Aug 27

StoSignSGD: Unbiased Structural Stochasticity Fixes SignSGD for Training Large Language Models

StoSignSGD is a new sign‑based optimization algorithm that injects structural stochasticity into the sign operator, ensuring unbiased updates. It resolves the divergence issues of traditional SignSGD on non‑smooth objectives, achieving optimal convergence rates in convex settings and improved complexity bounds in non‑convex, non‑smooth problems. Empirical results show that StoSignSGD is stable and efficient across large language model training, outperforming AdamW and SignSGD in low‑precision regimes (FP8 and FP4) and delivering speedups and accuracy gains on models ranging from OLMo2‑370M to 7B LLMs.

By Dingzhi Yu, Rui Pan, Yuxing Liu, Difan Zou, Tong Zhang
arXiv AI
Jun 4

Spectral Scaling Laws of Muon

arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.

By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar