arXiv Machine Learning

An Inertial Block Proximal Linearized Method with Adaptive Momentum for Nonconvex and Nonsmooth Optimization

arXiv:2608. 05502v1 Announce Type: cross Abstract: In this paper, we consider a class of multiblock nonconvex nonsmooth optimization problems, which covers many applications such as the analysis of pre-earthquake anomalies and machine learning.

arXiv Machine Learning
5d ago

To Solve Bilevel Optimization with Nonconvex Lower Levels, We Need Second-Order Stationarity

arXiv:2609. 30501v1 Announce Type: new Abstract: Although bilevel optimization (BLO) has emerged as a powerful framework for addressing many complex and nested machine learning problems in recent years, most existing studies are confined to the lower-level strongly convex (LLSC) or lower-level generally convex (LLGC) settings (i.

By Zhiyao Zhang, Menglu Yu, Alvaro Velasquez, Nathaniel D. Bastian, Jia Liu
arXiv Machine Learning
Aug 28

A unified convergence theory for adaptive first-order methods in the nonconvex case, including AdaNorm, full and diagonal AdaGrad and Muon

The paper introduces a unified framework for first‑order optimization algorithms applied to nonconvex unconstrained problems. It incorporates adaptively preconditioned gradients and covers popular methods such as full and diagonal AdaGrad, AdaNorm, and an adaptive variant of Muon. The framework supports heterogeneous geometries across variable groups and provides a fully stochastic global convergence analysis for all methods, with or without two types of momentum, under reasonable variance assumptions without requiring bounded stochastic gradients or small step sizes.

By S. Gratton, Ph. L. Toint