arXiv:2606. 28486v1 Announce Type: cross Abstract: The emergence of low-dimensional structures in the spectra of neural network weight matrices is a common empirical feature of trained models, but the dynamical origin of this phenomenon during learning remains an open problem.
By Chanju Park, Dario Bocchi, Francesco D'Amico, Biagio Lucini, Gert Aarts
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat
arXiv:2311. 02960v5 Announce Type: replace Abstract: Over the past decade, deep learning has proven to be a highly effective tool for learning meaningful features from raw data.
By Peng Wang, Xiao Li, Can Yaras, Zhihui Zhu, Laura Balzano, Wei Hu, Qing Qu
arXiv:2607. 21366v1 Announce Type: cross Abstract: Deep neural networks encode complex representations, but deconstructing this internal knowledge remains a challenge.
By Hossein Mobahi, Peter L. Bartlett
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.