arXiv:2608. 12009v1 Announce Type: cross Abstract: Bregman proximal stochastic gradient (BPSG) methods bring variance-reduced composite optimization to objectives whose geometry is poorly captured by Euclidean smoothness.
By Chenhan Jin, Shengze Xu, Binghui Xie, Kaiwen Zhou, Fan Jia, James Cheng, Tieyong Zeng
The paper introduces a novel technique called "persistence of memory" to enhance stochastic subspace methods for large‑scale optimisation. By using a weakly correlated guidance vector that is refreshed only at wide intervals, the method provides a structured direction for random subspace descent. The authors demonstrate that this guidance can be efficiently computed in sparse or minibatch settings and present the first theoretical analysis of classical SSD methods for sparse functions, showing alignment with low‑lying Hessian eigenvectors near the optimum.
By Subhroshekhar Ghosh, Clement Z. Q. Ng, Pierre-Louis Poirion, Akiko Takeda
arXiv:2510. 01878v2 Announce Type: replace Abstract: Low-rank gradient optimization for large language models is currently divided into two categories: structured methods that rigorously identify subspaces, and randomized approaches employed primarily for computational efficiency.
By Sahar Rajabi, Nayeema Nonta, Sirisha Rambhatla
arXiv:2606. 12120v1 Announce Type: new Abstract: Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature.
By Pratik Jawanpuria, Bamdev Mishra
arXiv:2607. 21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning.
By Zhuan Liang, Zheng Zhai
arXiv:2606. 23867v1 Announce Type: new Abstract: The exact computation of the Normalized Maximum Likelihood (NML) codelength for regular non-smooth estimators (e.
By Trenton Lau, Gary P. T. Choi
The paper introduces a Projected Riemannian Gradient Descent (RGD) algorithm for computing the Bures‑Wasserstein barycenter of positive definite matrices, achieving dimension‑independent linear convergence at unit step size. It resolves a previous dichotomy by showing that clipping eigenvalues to a fixed interval yields a closed‑form, non‑expansive projection in the BW metric, allowing the algorithm to match the empirical speed of unit‑step RGD while maintaining theoretical guarantees. The method also extends to the invariant matrix projection problem, providing a unified dimension‑independent analysis.
arXiv:2607. 25299v1 Announce Type: cross Abstract: Optimization over the Stiefel manifold plays a significant role in various machine learning tasks.
By Yuan Zhang, Jiang Hu, Zhijian Lai, Lin Lin, Zaiwen Wen
arXiv:2608. 08642v1 Announce Type: new Abstract: We study exact Kullback--Leibler (KL) projection for low-rank factorizations whose two nonnegative factors have prescribed row marginals and a shared, learned column marginal.
By Enliang Hu
The paper introduces a learning-based surrogate approach for stochastic optimization problems where uncertainty depends on the decision, modeled via a nonparametric regression. It constructs a surrogate that embeds iteratively updated Jacobian estimates, using an adaptive random design that focuses sampling near the current iterate to achieve dimension‑independent convergence of the Jacobian estimates. The resulting learning‑based stochastic prox‑linear (L‑SPL) algorithm demonstrates nonasymptotic convergence rates and outperforms existing methods in sample efficiency and objective value in numerical experiments.
By Boyang Shen, Junyi Liu
arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.
By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv:2607. 08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian.
By Lachlan Ewen MacDonald, Ren\'e Vidal