arXiv Machine Learning

Establishing Boundary KKT Convergence of Mirror Descent through Reparameterization

arXiv:2608. 07248v1 Announce Type: cross Abstract: We prove that mirror descent converges to a KKT point for the nonconvex problem without excluding boundary limits.

arXiv Machine Learning
Aug 4

Non-KKT Accumulation in Entropic Mirror Descent

arXiv:2608. 01658v1 Announce Type: cross Abstract: For mirror descent generated by a Legendre kernel, perhaps one of the most basic question in optimization is this: must every accumulation point of a bounded mirror descent sequence be Karush--Kuhn--Tucker (KKT) stationary under proper stepsizes?

By Kuangyu Ding, Kim-Chuan Toh
arXiv Statistics ML
Sep 17

Fenchel-Young Duality Gaps: Certified Early Stopping for Regularized Inverse Problems

The paper introduces computable error bounds and a certified early‑stopping criterion for regularized inverse problems by exploiting an exact Fenchel–Young duality‑gap identity. The total duality gap splits into a data‑fidelity loss and a regularizer loss, both expressed as Fenchel–Young losses that are oracle‑free and vanish exactly at Mirror Alignment. Using a constructive Brønsted–Rockafellar approach, the authors build a dual‑feasible proxy via a proximal step in the fidelity geometry, enabling an early‑stopping rule based on the regularizer loss.

By Pierre-Cyril Aubin-Frankowski (CERMICS UMR 9032, ENPC), Yohann de Castro (ICJ, ECL, IUF, PSPM)
arXiv Machine Learning
Jul 22

Linear convergence of proximal descent schemes on the Wasserstein space

arXiv:2411. 15067v2 Announce Type: replace-cross Abstract: We investigate proximal descent methods, inspired by the minimizing movement scheme introduced by Jordan, Kinderlehrer and Otto, for optimizing entropy-regularized functionals on the Wasserstein space.

By Razvan-Andrei Lascu, Mateusz B. Majka, David \v{S}i\v{s}ka, {\L}ukasz Szpruch
Hugging Face Trending Papers
Sep 3

Projected Riemannian Gradient Descent for the Bures-Wasserstein Barycenter: Dimension-Independent Linear Convergence at Unit Step Size

The paper introduces a Projected Riemannian Gradient Descent (RGD) algorithm for computing the Bures‑Wasserstein barycenter of positive definite matrices, achieving dimension‑independent linear convergence at unit step size. It resolves a previous dichotomy by showing that clipping eigenvalues to a fixed interval yields a closed‑form, non‑expansive projection in the BW metric, allowing the algorithm to match the empirical speed of unit‑step RGD while maintaining theoretical guarantees. The method also extends to the invariant matrix projection problem, providing a unified dimension‑independent analysis.

arXiv Machine Learning
Jun 11

Mirror Descent Beyond Euclidean Stability: An Exponential Separation in Initialization Sensitivity

arXiv:2606. 11431v1 Announce Type: new Abstract: Mirror Descent (MD) extends Gradient Descent (GD) beyond Euclidean geometry and has recently reappeared as a lens for KL-regularized policy optimization in reinforcement learning and LLM post-training.

By Shira Vansover-Hager, Matan Schliserman, Ofir Schlisselberg, Tomer Koren
arXiv Machine Learning
4d ago

Averaged Mirror Descent and Dual Gradient Methods: Convergent Algorithms for Entropic Gromov-Wasserstein Problems

The paper studies algorithms for computing the Entropic Gromov-Wasserstein (EGW) distance, a measure of discrepancy between metric measure spaces. It introduces Averaged Mirror Descent (AMD), which averages successive Mirror Descent steps and is proven to converge for any cost function, and shows that a dual gradient method with a fixed step size also converges for arbitrary costs, even when iterations are inexact. Empirical comparisons demonstrate that both AMD and the dual gradient method succeed on cases where classical Mirror Descent fails.

By Joanna Marks, Gabriel Rioux, Riccardo Passeggeri
arXiv Machine Learning
Jul 21

Scaling Limits of Constant-Stepsize SGD at Flat Minima

arXiv:2607. 16384v1 Announce Type: new Abstract: For stochastic gradient descent (SGD) with a constant stepsize $\alpha$, the invariant law of the iterates, centered at a minimizer, describes the behavior of the algorithm over long time horizons.

By Jingyi Zhang, Cheng Mao, Debankur Mukherjee