arXiv:2502. 15952v3 Announce Type: replace Abstract: Recent works exploring the training dynamics of homogeneous neural network weights under gradient flow with small initialization have established that in the early stages of training, the weights remain small and near the origin, but converge in direction.
By Akshay Kumar, Jarvis Haupt
arXiv:2606. 04327v1 Announce Type: cross Abstract: We investigate the geometric structure of stationary plateaus that arise in the loss landscape of two-layer neural networks with smooth activation functions.
By Tian Ding, Dawei Li, Ruoyu Sun
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat
The paper investigates how two‑layer polynomial‑width neural networks learn orthogonal multi‑index targets under standard initialization. It shows that incremental learning still occurs: the loss decreases sequentially following the Hermite expansion, with lower‑order components learned first. The dynamics also exhibit a competitive reallocation of parameter mass, shifting into the target subspace and concentrating on aligned neurons. The analysis uses a symmetry‑based finite‑width approximation and demonstrates that vanilla gradient descent displays the same qualitative behavior.
By Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel
arXiv:2402. 00152v5 Announce Type: replace Abstract: Constructing the architecture of a neural network is a challenging pursuit for the machine learning community, and the dilemma of whether to go deeper or wider remains a persistent question.
By Yahong Yang, Juncai He
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi