arXiv AI By Wenpeng Zhang, Runsheng Yu, Peilin Zhao

Matrix AdaGrad: Row-wise and Column-wise Adaptive Subgradient Methods

Read the original on arXiv AI →

The paper introduces Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad, two adaptive subgradient methods that extend AdaGrad to matrix-valued parameters by using row-wise and column-wise proximal functions. It presents a general Online Mirror Descent framework that derives these optimizers through online regret minimization, providing regret guarantees that can be tighter than entry-wise AdaGrad for structured gradients. Experiments on matrix factorization and deep neural-network training show that aligning adaptive scaling with matrix structure improves optimization stability, allows larger learning rates, and supports greater network depth.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Jul 31

Towards Stability of Parameter-Free Optimization

arXiv:2405. 04376v4 Announce Type: replace Abstract: Hyperparameter tuning, particularly the selection of an appropriate learning rate in adaptive gradient training methods, remains a challenge.

By Yijiang Pang, Shuyang Yu, Bao Hoang, Jiayu Zhou