arXiv:2607. 07204v2 Announce Type: replace-cross Abstract: Structured preconditioners restrict optimization to a small family of positive metrics, but endpoint condition-number reachability does not measure the geometric effort required to reach a useful metric.
By Zavier Li
arXiv:2607. 06723v1 Announce Type: cross Abstract: Most gradient-based optimization methods move parameters through a fixed background geometry, even when their internal states implicitly define changing notions of length, curvature, and preconditioning.
By Zavier Li
arXiv:2607. 25624v1 Announce Type: new Abstract: Positive quadratic networks admit the low-rank representation f_U(x)=x^top UU^top x, where Uinmathbb{R}^{dtimes r} is identifiable only up to right orthogonal multiplication, representing a rank-r PSD matrix Q=UU^top.
By Pengcheng Cheng
arXiv:2607. 07206v1 Announce Type: new Abstract: Adaptive optimizers mix several mechanisms: a metric or preconditioner maps gradients to descent directions, while estimation, memory, step-size control, constraints, stochasticity, target modification, and discretization determine which directions are available and how they are used.
By Zavier Li
arXiv:2209. 15130v3 Announce Type: replace-cross Abstract: We study a general matrix optimization problem with a fixed-rank positive semidefinite (PSD) constraint.
By Yuetian Luo, Nicolas Garcia Trillos
arXiv:2607. 07204v1 Announce Type: cross Abstract: Optimization geometrodynamics views optimizer state as evolving geometry.
By Zavier Li
arXiv:2607. 08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian.
By Lachlan Ewen MacDonald, Ren\'e Vidal
arXiv:2601. 21487v2 Announce Type: replace-cross Abstract: We study minimization of smooth functions over feasible sets that have smooth embedded-manifold structure throughout or only on selected regions, using linear minimization oracles (LMOs) to determine search directions under user-chosen norms.
By Kaiwei Yang, Lexiao Lai
arXiv:2608. 13201v1 Announce Type: cross Abstract: We develop the statistical and algorithmic theory of inverse optimal transport (IOT) under the feature-parameterized cost C_theta(i,j) = -theta^T phi(i,j).
By Han Dong, Jiaming Li, Yongqiang Gong, Ruixi Li, Yin Liu
arXiv:2606. 13825v1 Announce Type: cross Abstract: Deep unfolding (DU) accelerates iterative optimizers by introducing learnable components and training them through unrolled iterations, but extending DU to the large-scale semidefinite programs (SDPs) common in robotics has remained limited.
By Alex Oshin, Rahul Vodeb Ghosh, Evangelos A. Theodorou
arXiv:2607. 23642v1 Announce Type: cross Abstract: Discrete optimization algorithms are often analyzed through continuous-time limiting ODEs, but a convergence certificate for the ODE is not automatically one for the discrete algorithm.
By George A Kevrekidis
arXiv:2607. 22004v1 Announce Type: new Abstract: Energy natural gradient descent (ENGD) aligns parameter updates with the curvature of an underlying function-space energy, but existing formulations assume an unconstrained Euclidean parameter domain.
By Zhangyong Liang, Huanhuan Gao