arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
By Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik, Pierre Youssef
arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.
By Ori Shem-Ur, Yaron Oz
The paper investigates how deep residual networks behave when their initial weights are correlated across layers. It confirms a conjecture that such correlated initializations interpolate between a Brownian stochastic differential equation (for independent weights) and an ordinary differential equation (for perfectly correlated weights). By applying a feature function to a stationary Gaussian sequence with regularly varying correlation, the authors prove that a unique critical scaling exists, leading the infinite‑depth limit to a Young differential equation driven by a Hermite process, which reduces to fractional Brownian motion when the feature function has Hermite rank one. The study shows that the correlation structure and Hermite rank of the initialization uniquely determine the critical scaling and asymptotic limit, making them meaningful hyperparameters in the asymptotic regime, whereas finite‑variance i.i.d. initialization always yields a Brownian driver regardless of distribution.
By Felix Benning, Ivan Nourdin, Giovanni Peccati
arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
By Constantin Kogler, Tassilo Schwarz, Samuel Kittle
arXiv:2607. 12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization.
By Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.
arXiv:2508. 11522v4 Announce Type: replace Abstract: Neural tangent kernels (NTKs) are a powerful tool for analyzing deep, non-linear neural networks.
By Max Guillen, Philipp Misof, Jan E. Gerken
arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.
By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna
The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher...
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2607. 07884v1 Announce Type: new Abstract: In this short note we consider the gradient descent dynamics of deep scalar linear networks, $f(x) = \prod_{l=1}^L w_l x$, which enjoy exact time-course solutions for any integer depth.
By Yedi Zhang, Peter E. Latham, Leena Chennuru Vankadara, Andrew Saxe
arXiv:2607. 14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization.
By Katie Everett