arXiv Machine Learning

Riemannian Gradient Descent for Low-Rank Architectures

arXiv:2606. 02328v1 Announce Type: new Abstract: We explore Riemannian optimization techniques for rank-factored matrix parameters, targeting contemporary deep learning applications.

arXiv AI
Sep 24

Riemannian Structure and Optimization for a Class of Low-Parametric Orthogonal Matrices

The paper studies matrices built from block‑diagonal factors interleaved with fixed permutations, a structured family useful in deep learning for balancing expressivity and efficiency. By applying Riemannian geometry, the authors determine when this class forms a smooth manifold and develop Riemannian tools for the orthogonal two‑factor case. They propose efficient algorithms that use automatic differentiation, allow parameter sharing, and avoid dense matrix construction, testing them on matrix approximation and fine‑tuning large language models, while also exploring properties of factorizations with more factors.

By Ali Aliev, Maxim Rakhuba
Hugging Face Trending Papers
Jul 14

Reassessing Muon for Matrix Factorization

Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.

arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv Machine Learning
Aug 18

SubZero+: Efficient Zeroth-Order LLM Fine-Tuning via Large Learning Rates

arXiv:2608. 15665v1 Announce Type: new Abstract: Zeroth-order (ZO) optimization enables backpropagation-free fine-tuning of large language models, but existing ZO methods suffer from high-variance gradient estimators, making convergence unstable and highly sensitive to learning rates.

By Ziming Yu, Shuyao Xiao, Xingyu Zhao, Sike Wang, Pan Zhou, Peiyu Zang, Xiangda Yan, Yongjie Yang, Jia Li
arXiv Machine Learning
Jul 17

Stabilizing Native Low-Rank LLM Pretraining

arXiv:2602. 12429v2 Announce Type: replace Abstract: Foundation models have achieved remarkable success, yet their growing parameter counts pose significant computational and memory challenges.

By Paul Janson, Edouard Oyallon, Eugene Belilovsky
arXiv Machine Learning
Jun 2

Riemannian Optimization for Hadamard Products of Low-Rank Matrices

arXiv:2606. 01216v1 Announce Type: new Abstract: The elementwise Hadamard product of two low-rank matrices provides a parameter-efficient model for data with multiplicative structure, but its modeling is challenging due to the presence of additional symmetries under coupled row/column scalings between the two factors.

By Pratik Jawanpuria, Ankish Chandresh, Bamdev Mishra
arXiv Machine Learning
Sep 23

MONA: Muon Optimizer with Nesterov Acceleration for Scalable Language Model Training

MONA is a new optimizer that extends the Muon optimizer by adding a Nesterov‑style acceleration term derived from an exponential moving average of gradient differences. The paper provides a convergence analysis showing that this term offers curvature‑aware corrections while maintaining Muon’s spectral‑norm regularization. Empirical results demonstrate that MONA outperforms both Muon and AdamW on Mixture‑of‑Experts pretraining across models ranging from 1 B to 68 B parameters, and achieves state‑of‑the‑art performance on downstream benchmarks after fine‑tuning the largest model.

By Jiacheng Li, Jianchao Tan, Hongtao Xu, Jiaqi Zhang, Yifan Lu, Yerui Sun, Yuchen Xie, Xunliang Cai
Hugging Face Trending Papers
Sep 24

Automatic Rank Allocation for Low-Rank Adaptation in Large Language Models via lp Regularization

The paper introduces αp-LoRA, a rank-allocation strategy for low-rank adaptation (LoRA) in large language models that uses π-regularization (0 < p < 1) to induce sparsity in rank-one components. By regularizing the energy of each component, redundant parts are encouraged to vanish while important ones are retained, and the authors derive a proximal subproblem that reduces the matrix optimization to a two‑dimensional thresholding criterion. Experiments on natural language understanding and question‑answering tasks show that αp-LoRA achieves performance competitive with existing LoRA baselines.