arXiv:2506. 13139v3 Announce Type: replace-cross Abstract: Modern Machine Learning (ML) and Deep Neural Networks (DNNs) often operate on high-dimensional data and rely on overparameterized models, where classical low-dimensional intuitions break down.
By Zhenyu Liao, Michael W. Mahoney
arXiv:2403. 10232v2 Announce Type: replace-cross Abstract: Conventional matrix completion methods approximate the missing values by assuming the matrix to be low-rank, which leads to a linear approximation of missing values.
By Sajad Faramarzi, Farzan Haddadi, Sajjad Amini, Masoud Ahookhosh, Symeon Chatzinotas
arXiv:2607. 08783v1 Announce Type: cross Abstract: Manifold-valued measurements are prevalent in various machine learning tasks.
By Ziheng Chen, Yue Song, Rui Wang, Xiao-Jun Wu, Nicu Sebe
arXiv:2607. 07735v1 Announce Type: cross Abstract: Sparse precision matrix estimation provides an interpretable and computationally efficient framework for modeling conditional dependencies in high-dimensional, low-sample-size data.
By Aryan Eftekhari, Daniel Sergio Vega, Ernst-Jan Camiel Wit, Olaf Schenk
arXiv:2606. 31390v1 Announce Type: cross Abstract: Low-rank matrix optimization is often carried out via the Burer-Monteiro (BM) formulation, but choosing the factorization rank $r$ is delicate and can substantially slow optimization.
By Yudong Wei, Liang Zhang, Bingcong Li, Niao He
arXiv:2607. 13246v1 Announce Type: cross Abstract: Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training.
By Ali Parviz, Gal Mishne, Alex Cloninger
arXiv:2607. 11938v1 Announce Type: cross Abstract: This book is about the mathematical foundations of data science.
By Afonso S. Bandeira, Amit Singer, Thomas Strohmer
arXiv:2606. 15702v1 Announce Type: cross Abstract: Modern deep learning optimization features heterogeneous parameter structures, noisy gradients, and highly nonconvex landscapes, posing significant challenges for both algorithm design and theoretical analysis.
By Bohao Ma, Junyu Zhang, Chuan He
Muon has recently emerged as a strong optimizer for large-scale deep learning, where it reshapes gradient updates through approximate orthogonalization and has been reported to outperform Adam and AdamW in large language model training. Its empirical success has motivated a growing body of theoretical work that interprets Muon as steepest descent under the spectral norm.
arXiv:2509. 24882v2 Announce Type: replace Abstract: Neural scaling laws underlie many of the recent advances in deep learning, yet their theoretical understanding remains largely confined to linear models.
By Leonardo Defilippis, Yizhou Xu, Julius Girardin, Emanuele Troiani, Vittorio Erba, Lenka Zdeborov\'a, Bruno Loureiro, Florent Krzakala
arXiv:2503. 11891v2 Announce Type: replace Abstract: We analyze the landscape and training dynamics of diagonal linear networks in a linear regression task, with the network parameters being perturbed by isotropic normal noise during training.
By Gabriel Clara, Sophie Langer, Johannes Schmidt-Hieber