arXiv Machine Learning By S. Gratton, Ph. L. Toint

A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad and Muon

Read the original on arXiv Machine Learning →

The paper introduces a unified framework for first‑order optimization algorithms applied to nonconvex unconstrained problems. It incorporates adaptively preconditioned gradients and covers popular methods such as full and diagonal AdaGrad, AdaNorm, and an adaptive variant of Muon. The framework supports heterogeneous geometries across variable groups and provides a fully stochastic global convergence analysis for all methods, with or without two types of momentum, under reasonable variance assumptions without requiring bounded stochastic gradients or small step sizes.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 18

Adaptive Optimization via Momentum on Variance-Normalized Gradients

arXiv:2602. 10204v2 Announce Type: replace Abstract: We introduce MVN-Grad (Momentum on Variance-Normalized Gradients), an Adam-style optimizer that improves stability and performance by combining two complementary ideas: variance-based normalization and momentum applied after normalization.

By Francisco Patitucci, Aryan Mokhtari