arXiv:2607. 07206v1 Announce Type: new Abstract: Adaptive optimizers mix several mechanisms: a metric or preconditioner maps gradients to descent directions, while estimation, memory, step-size control, constraints, stochasticity, target modification, and discretization determine which directions are available and how they are used.
By Zavier Li
arXiv:2607. 07204v1 Announce Type: cross Abstract: Optimization geometrodynamics views optimizer state as evolving geometry.
By Zavier Li
arXiv:2607. 07204v2 Announce Type: replace-cross Abstract: Structured preconditioners restrict optimization to a small family of positive metrics, but endpoint condition-number reachability does not measure the geometric effort required to reach a useful metric.
By Zavier Li
arXiv:2510. 01943v3 Announce Type: replace-cross Abstract: Quasar-convex functions form a broad nonconvex class with applications to linear dynamical systems, generalized linear models, and Riemannian optimization, among others.
By David Mart\'inez-Rubio
The paper introduces a new convergence framework for solving distributionally robust optimization problems formulated as nonconvex, nonconcave minimax problems over a Euclidean space and a Riemannian manifold. It defines a "basin saddle point"—a locally defined Nash equilibrium—and proves that a Riemannian gradient ascent–descent algorithm converges to such points under a local Łojasiewicz growth condition. The authors apply this theory to a statistical risk DRO problem over Gaussian measures, deriving explicit convergence rates and constants in terms of data dimension, loss moments, and reference covariance.
By Rishabh Dixit, Pranav Upadrashta, Alex Cloninger
arXiv:2607. 22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain.
By Zhangyong Liang, Huanhuan Gao
arXiv:2607. 06723v2 Announce Type: replace-cross Abstract: Adaptive optimizers carry hidden states that change how visible gradients become parameter motion.
By Zavier Li
arXiv:2606. 12120v1 Announce Type: new Abstract: Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature.
By Pratik Jawanpuria, Bamdev Mishra
arXiv:2608. 02487v1 Announce Type: cross Abstract: Recently, rectified flow has emerged as a fundamental framework for large-scale image generation, powering state-of-the-art systems such as FLUX.
By Leda Wang, Zhehao Xu, Qiang Liu, Harrison H. Zhou
arXiv:2609. 17089v1 Announce Type: cross Abstract: The choice of Riemannian metric can strongly influence the convergence of gradient-based optimization over covariance matrices.
By Yibang Li, Bamdev Mishra, Pratik Jawanpuria, Cyrus Mostajeran
arXiv:2510. 21033v3 Announce Type: replace-cross Abstract: We develop a theory of iso-Riemannian optimization for problems constrained to learned data manifolds, a setting in which classical Riemannian optimization - and Riemannian gradient descent in particular - can be poorly suited.
By Willem Diepeveen, Melanie Weber
The limited-memory BFGS (L-BFGS) algorithm is a cornerstone of large-scale optimization due to its linear memory and computational costs. However, in ill-conditioned or non-convex landscapes, the implicit inverse Hessian approximation can suffer from an exploding condition number, leading to numerical instability and degraded convergence.