arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
By Zhenyu Liao, Michael W. Mahoney
The paper introduces a deterministic Simulated Oracle Direction (SOD) framework that enables escaping spurious local minima in non‑convex low‑rank matrix sensing without explicit tensor lifting. By projecting over‑parameterized escape directions back into the original parameter space, the method guarantees a strict decrease in objective value from existing local minima. Experiments show reliable escape from local minima and convergence to global optima with minimal computational overhead compared to explicit over‑parameterization.
By Tianqi Shen, Jinji Yang, Junze He, Kunhan Gao, Zeyu Zheng, Ziye Ma
arXiv:2403. 10232v2 Announce Type: replace-cross Abstract: Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values.
By Sajad Faramarzi, Farzan Haddadi, Sajjad Amini, Masoud Ahookhosh, Symeon Chatzinotas
arXiv:2610.02182v1 Announce Type: cross
Abstract: Quasi-Newton (QN) methods have long been among the most effective methods for large-scale unconstrained convex optimization. Two obstacles have limit...
By Joohwan Ko, Tetiana Parshakova, Diana Cai, Robert M. Gower
arXiv:2608.19021v2 Announce Type: replace
Abstract: Global Covariance Pooling (GCP) improves deep networks by capturing second-order feature statistics, and is especially effective for fine-grained r...
By Md Rifat Ur Rahman, Md Raihan Khan, Md Sakib Hossain Shovon, Pietro Li\`o, Mohammad Ali Moni
The paper introduces Row-wise Matrix AdaGrad and Column-wise Matrix AdaGrad, two adaptive subgradient methods that extend AdaGrad to matrix-valued parameters by using row-wise and column-wise proximal functions. It presents a general Online Mirror Descent framework that derives these optimizers through online regret minimization, providing regret guarantees that can be tighter than entry-wise AdaGrad for structured gradients. Experiments on matrix factorization and deep neural-network training show that aligning adaptive scaling with matrix structure improves optimization stability, allows larger learning rates, and supports greater network depth.
By Wenpeng Zhang, Runsheng Yu, Peilin Zhao