arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
By Zhenyu Liao, Michael W. Mahoney
The paper introduces a deterministic Simulated Oracle Direction (SOD) framework that enables escaping spurious local minima in non‑convex low‑rank matrix sensing without explicit tensor lifting. By projecting over‑parameterized escape directions back into the original parameter space, the method guarantees a strict decrease in objective value from existing local minima. Experiments show reliable escape from local minima and convergence to global optima with minimal computational overhead compared to explicit over‑parameterization.
By Tianqi Shen, Jinji Yang, Junze He, Kunhan Gao, Zeyu Zheng, Ziye Ma
arXiv:2403. 10232v2 Announce Type: replace-cross Abstract: Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values.
By Sajad Faramarzi, Farzan Haddadi, Sajjad Amini, Masoud Ahookhosh, Symeon Chatzinotas
arXiv:2610.02182v1 Announce Type: cross
Abstract: Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limit...
By Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower
arXiv:2608.19021v2 Announce Type: replace
Abstract: Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained r...
By Md Rifat Ur Rahman, Md Raihan Khan, Md Sakib Hossain Shovon, Pietro Li\`o, Mohammad Ali Moni
The paper introduces Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad, two adaptive subgradient methods that extend AdaGrad to matrix-valued parameters by using row-wise and column-wise proximal functions. It presents a general Online Mirror Descent framework that derives these optimizers through online regret minimization, providing regret guarantees that can be tighter than entry-wise AdaGrad for structured gradients. Experiments on matrix factorization and deep neural-network training show that aligning adaptive scaling with matrix structure improves optimization stability, allows larger learning rates, and supports greater network depth.
By Wenpeng Zhang, Runsheng Yu, Peilin Zhao
arXiv:2607. 08783v1 Announce Type: cross Abstract: Manifold-valued measurements are prevalent in various machine learning tasks.
By Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, Nicu Sebe
arXiv:2607. 07735v1 Announce Type: cross Abstract: Sparse precision matrix estimation provides an interpretable and computationally efficient framework for modeling conditional dependencies in high-dimensional, low-sample-size data.
By Aryan Eftekhari, Daniel Sergio Vega, Ernst-Jan Camiel Wit, Olaf Schenk
arXiv:2606. 31390v1 Announce Type: cross Abstract: Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank $r$ is delicate and can substantially slow optimization.
By Yudong Wei, Liang Zhang, Bingcong Li, Niao He
arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.
By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv:2607. 11938v1 Announce Type: cross Abstract: This book is about the mathematical foundations of data science.
By Afonso S. Bandeira, Amit Singer, Thomas Strohmer
arXiv:2606. 15702v1 Announce Type: cross Abstract: Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis.
By Bohao Ma, Junyu Zhang, Chuan He