arXiv Machine Learning

On the Benefits of Weight Normalization for Overparameterized Matrix Sensing

arXiv:2510. 01175v2 Announce Type: replace Abstract: While normalization techniques are widely used in deep learning, their theoretical understanding remains relatively limited.

arXiv Machine Learning
Aug 20

Escaping Local Minima Provably in Non-convex Matrix Sensing: A Deterministic Framework via Simulated Lifting

The paper introduces a deterministic Simulated Oracle Direction (SOD) framework that enables escaping spurious local minima in non‑convex low‑rank matrix sensing without explicit tensor lifting. By projecting over‑parameterized escape directions back into the original parameter space, the method guarantees a strict decrease in objective value from existing local minima. Experiments show reliable escape from local minima and convergence to global optima with minimal computational overhead compared to explicit over‑parameterization.

By Tianqi Shen, Jinji Yang, Junze He, Kunhan Gao, Zeyu Zheng, Ziye Ma
arXiv AI
Sep 21

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

The paper introduces Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad, two adaptive subgradient methods that extend AdaGrad to matrix-valued parameters by using row-wise and column-wise proximal functions. It presents a general Online Mirror Descent framework that derives these optimizers through online regret minimization, providing regret guarantees that can be tighter than entry-wise AdaGrad for structured gradients. Experiments on matrix factorization and deep neural-network training show that aligning adaptive scaling with matrix structure improves optimization stability, allows larger learning rates, and supports greater network depth.

By Wenpeng Zhang, Runsheng Yu, Peilin Zhao
arXiv AI
Jul 16

Reassessing Muon for Matrix Factorization

arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.

By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv Machine Learning
Jun 16

Schattor: Schatten-family methods for deep learning optimization

arXiv:2606. 15702v1 Announce Type: cross Abstract: Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis.

By Bohao Ma, Junyu Zhang, Chuan He