arXiv:2606. 04327v1 Announce Type: cross Abstract: We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions.
By Tian Ding, Dawei Li, Ruoyu Sun
arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
By Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
arXiv:2607. 23397v1 Announce Type: new Abstract: Hierarchical neural networks are widely used in artificial intelligence, yet their mathematical properties remain incompletely understood.
By Sumio Watanabe
arXiv:2608. 06766v1 Announce Type: cross Abstract: Training changes a network's predictions while allocating task-relevant structure across its internal units.
By Tongxi Wang
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat
arXiv:2607. 12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization.
By Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2507. 14177v2 Announce Type: replace-cross Abstract: This paper aims to understand the training solution, which is obtained by the back-propagation algorithm, of two-layer neural networks whose hidden layer is composed of the units with smooth activation functions, including the usual sigmoid type most commonly used before the advent of ReLUs.
By Changcun Huang
arXiv:2507. 05164v2 Announce Type: replace-cross Abstract: In this chapter, we utilize dynamical systems to analyze several aspects of machine learning algorithms.
By Dennis Chemnitz, Maximilian Engel, Christian Kuehn, Sara-Viola Kuntz
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.
arXiv:2606. 28464v1 Announce Type: new Abstract: In the optimization of neural networks, gradient dynamics are influenced by critical points that arise from the model's architecture.
By Kathl\'en Kohn, Giovanni Luca Marchetti, Farhan Shabir, Vahid Shahverdi, Weisheng Wang