Critical initialization destabilizes higher input derivatives in wide scalar-input networks
Read the original on arXiv Statistics ML →The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher...
arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.
arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
arXiv:2608. 14638v1 Announce Type: new Abstract: In this paper we study autoencoders, a special class of deep neural nets (DNNs) whose performance can be characterized via their fixed points.
arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.
The paper investigates how deep residual networks behave when their initial weights are correlated across layers. It confirms a conjecture that such correlated initializations interpolate between a Brownian stochastic differential equation (for independent weights) and an ordinary differential equation (for perfectly correlated weights). By applying a feature function to a stationary Gaussian sequence with regularly varying correlation, the authors prove that a unique critical scaling exists, leading the infinite‑depth limit to a Young differential equation driven by a Hermite process, which reduces to fractional Brownian motion when the feature function has Hermite rank one. The study shows that the correlation structure and Hermite rank of the initialization uniquely determine the critical scaling and asymptotic limit, making them meaningful hyperparameters in the asymptotic regime, whereas finite‑variance i.i.d. initialization always yields a Brownian driver regardless of distribution.