arXiv:2608. 12009v1 Announce Type: cross Abstract: Bregman proximal stochastic gradient (BPSG) methods bring variance-reduced composite optimization to objectives whose geometry is poorly captured by Euclidean smoothness.
By Chenhan Jin, Shengze Xu, Binghui Xie, Kaiwen Zhou, Fan Jia, James Cheng, Tieyong Zeng
arXiv:2510. 01878v2 Announce Type: replace Abstract: Low-rank gradient optimization for large language models is currently divided into two categories: structured methods that rigorously identify subspaces, and randomized approaches employed primarily for computational efficiency.
By Sahar Rajabi, Nayeema Nonta, Sirisha Rambhatla
arXiv:2606. 12120v1 Announce Type: new Abstract: Low-rank optimal transport (OT) mitigates the quadratic scaling of classical solvers, yet existing approaches rely heavily on first-order mirror-descent updates that require careful hyperparameter tuning and ignore the optimization landscape's curvature.
By Pratik Jawanpuria, Bamdev Mishra
arXiv:2607. 21039v1 Announce Type: new Abstract: Spectral methods are among the most widely used techniques for community detection, clustering, and graph learning.
By Zhuan Liang, Zheng Zhai
arXiv:2606. 23867v1 Announce Type: new Abstract: The exact computation of the Normalized Maximum Likelihood (NML) codelength for regular non-smooth estimators (e.
By Trenton Lau, Gary P. T. Choi
arXiv:2607. 25299v1 Announce Type: cross Abstract: Optimization over the Stiefel manifold plays a significant role in various machine learning tasks.
By Yuan Zhang, Jiang Hu, Zhijian Lai, Lin Lin, Zaiwen Wen
arXiv:2608. 08642v1 Announce Type: new Abstract: We study exact Kullback--Leibler (KL) projection for low-rank factorizations whose two nonnegative factors have prescribed row marginals and a shared, learned column marginal.
By Enliang Hu
arXiv:2607. 03871v1 Announce Type: new Abstract: Minimum maximum mean discrepancy (MMD) estimation has emerged as a robust and likelihood-free alternative to maximum likelihood estimation for parameter estimation.
By Sophia Seulkee Kang, Louis Sharrock, Xiaoyuan Cheng, Fran\c{c}ois-Xavier Briol, Zonghao Chen
arXiv:2607. 08380v1 Announce Type: new Abstract: An important quantity in the theory of gradient descent (GD) is the \emph{sharpness}, defined as the largest eigenvalue of the objective Hessian.
By Lachlan Ewen MacDonald, Ren\'e Vidal
arXiv:2606. 30455v1 Announce Type: new Abstract: The standard convergence analysis of mini-batch stochastic gradient descent (SGD) models gradient noise using a single variance term that treats all parameter directions equally, ignoring the fact that noise in high-curvature directions has less impact because learning rates are already constrained there.
By Muhammad Hamza (Indian Institute of Technology Kharagpur), Ayush Goel (Indian Institute of Technology Kharagpur)
arXiv:2606. 27298v1 Announce Type: cross Abstract: We study the fundamental problem of learning a high-dimensional Gaussian truncated to an unknown halfspace.
By Haitong Liu, Deepak Narayanan Sridharan, David Steurer, Manuel Wiedmer
arXiv:2608. 15121v1 Announce Type: cross Abstract: Sufficient dimension reduction (SDR) seeks the minimal subspace of the predictors that captures the full conditional distribution of the response, which is known as the central subspace (CS).
By Ye Tian