arXiv Machine Learning By Pengcheng Xie

Derivative-Free Structured Updates for Muon

Read the original on arXiv Machine Learning →

The paper introduces a derivative‑free framework for Muon‑style updates, replacing gradient‑based momentum with structured finite differences. Four variants—full entrywise recovery, random low‑rank surrogates, basis‑aligned rank‑one probing, and direct structured search—are explored, with basis‑aligned probing shown to be equivalent to coordinate finite differences up to scaling. Experiments on matrix regression, noisy‑gradient regression, a neural network, and a CartPole task demonstrate that random rank‑one probing can significantly reduce function evaluations, though at the expense of update accuracy, and that accurate function values can sometimes offset unreliable gradient oracles.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
Hugging Face Trending Papers
Jul 14

Reassessing Muon for Matrix Factorization

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.

arXiv Machine Learning
Sep 18

Low-rank Orthogonalization for Large-scale Matrix Optimization with Applications to Foundation Model Training

The paper introduces low‑rank orthogonalization, a technique that exploits the low‑rank nature of gradients in neural network training to perform matrix orthogonalization more efficiently. Building on this, the authors present low‑rank matrix‑signed gradient descent (MSGD) and a low‑rank variant of the Muon optimizer, showing through experiments that low‑rank Muon matches or surpasses vanilla Muon on GPT‑2 and LLaMA pretraining, especially for larger models. Theoretical analysis provides iteration‑complexity bounds for both low‑rank MSGD and low‑rank Muon under heavy‑tailed noise.

By Chuan He, Zhanwang Deng, Zhaosong Lu