arXiv Machine Learning

PoLoRA: A Preconditioned Orthogonalized LoRA Optimizer

arXiv:2607. 17620v1 Announce Type: new Abstract: Low-rank adaptation (LoRA) makes finetuning large language models cheaper by adding to each weight matrix a trainable low-rank update parameterized as the product of two matrices.

arXiv Machine Learning
Sep 14

Rank-Efficient LoRA via Joint Tangent-Space Optimization under Isotropic Curvature

The paper introduces ISO-LoRA, an optimizer that improves rank utilization in Low‑Rank Adaptation (LoRA) by coupling factor updates through spectral descent on the induced tangent perturbation in weight space. Experiments on GPT‑2 adaptation show that standard optimizers like AdamW concentrate updates in a few singular directions, whereas ISO-LoRA distributes energy more evenly, leading to higher effective rank and better downstream performance across 0.1B‑7B models. The authors provide theoretical guarantees under a stylized spiked‑gradient model and demonstrate that ISO-LoRA consistently outperforms factor‑wise optimizers, especially at moderate‑to‑large LoRA ranks.

By Zihan Zhu, Zhehang Du, Xuyang Chen, Tim Tsz-Kit Lau, Jiayuan Wu, X. Y. Han, Qi Long, Weijie Su
arXiv AI
Jun 12

LoRA-Muon: Spectral Steepest Descent on the Low-Rank Manifold

arXiv:2606. 12921v1 Announce Type: cross Abstract: Low-Rank Adaptation (LoRA) significantly reduces compute and memory costs for finetuning Deep Learning models but is often harder to tune than dense training: when using factor-wise optimizers such as AdamW, it is sensitive to initialization choices, its optimal learning rates transfer poorly across ranks, and it often fails to beat dense baselines.

By Franz Louis Cesista, Katherine Crowson, C\'edric Simal, Stella Biderman
arXiv Machine Learning
Sep 2

Variance-Adaptive Muon: Pre-Orthogonalization Variance Modulation for Efficient Language Model Pretraining

The paper introduces two variance‑adaptive variants of the Muon optimizer—Muon‑NSR and Muon‑VS—for language model pretraining. Both methods incorporate gradient‑variance information into Muon’s orthogonalization process without adding extra hyperparameters, preserving its spectral normalization structure. Experiments on Llama‑style and GPT‑2 models ranging from 125 M to 1.2 B parameters show that these variants outperform well‑tuned Muon baselines and achieve up to a 1.33× step‑to‑target speedup on Llama‑1.2B.

By Jingru Li, Yibo Fan, Huan Li
arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
Hugging Face Trending Papers
Sep 2

LoRA-TSD: Tangent-Space Spectral Descent for LoRA via Muon-Style Updates

Low-rank adaptation (LoRA) is the standard way to fine-tune large models, yet when its two factors are trained independently, the update ignores the geometry of the low-rank weight change it induces. We introduce LoRA-TSD, an optimizer that treats every LoRA step as a tangent vector of the fixed-rank matrix manifold and takes the spectral-norm steepest-descent step of Muon inside that tangent space, mapping the result back to the factors through a retraction native to the LoRA parametrization.

arXiv AI
Sep 10

FedSubMuon: Communication-Efficient Federated LLM Fine-Tuning via Structured Subspace Muon

FedSubMuon introduces a communication‑efficient federated fine‑tuning approach for large language models by optimizing compact coefficient matrices within shared structured subspaces, thereby keeping Muon’s matrix‑aware optimization while reducing client upload size. An accuracy‑oriented variant, FedSubMuon‑GT, further adapts subspace bases using projected gradients to better align with task‑relevant directions. Experiments on instruction tuning and mathematical reasoning demonstrate that FedSubMuon‑GT achieves the best overall accuracy on most dataset‑model pairs, while FedSubMuon outperforms all matched‑budget baselines and reduces communication by up to 5.5× on Llama‑1B and 1.4× on Qwen‑4B compared to the closest baseline.

By Shaolong Chen, Youming Tao, Shuzhen Chen, Falko Dressler, Qingqing Ye, Di Wang