arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
By Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik, Pierre Youssef
arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.
By Ori Shem-Ur, Yaron Oz
The paper investigates how deep residual networks behave when their initial weights are correlated across layers. It confirms a conjecture that such correlated initializations interpolate between a Brownian stochastic differential equation (for independent weights) and an ordinary differential equation (for perfectly correlated weights). By applying a feature function to a stationary Gaussian sequence with regularly varying correlation, the authors prove that a unique critical scaling exists, leading the infinite‑depth limit to a Young differential equation driven by a Hermite process, which reduces to fractional Brownian motion when the feature function has Hermite rank one. The study shows that the correlation structure and Hermite rank of the initialization uniquely determine the critical scaling and asymptotic limit, making them meaningful hyperparameters in the asymptotic regime, whereas finite‑variance i.i.d. initialization always yields a Brownian driver regardless of distribution.
By Felix Benning, Ivan Nourdin, Giovanni Peccati
arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
By Constantin Kogler, Tassilo Schwarz, Samuel Kittle
arXiv:2607. 12332v1 Announce Type: new Abstract: We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization.
By Jiajie Zhao, Jianxing Wang, Junjie Yang, Zhiwei Bai, Yaoyu Zhang
We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.