arXiv Machine Learning

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

arXiv Machine Learning
Sep 4

Correlated initialization of deep residual networks

The paper investigates how deep residual networks behave when their initial weights are correlated across layers. It confirms a conjecture that such correlated initializations interpolate between a Brownian stochastic differential equation (for independent weights) and an ordinary differential equation (for perfectly correlated weights). By applying a feature function to a stationary Gaussian sequence with regularly varying correlation, the authors prove that a unique critical scaling exists, leading the infinite‑depth limit to a Young differential equation driven by a Hermite process, which reduces to fractional Brownian motion when the feature function has Hermite rank one. The study shows that the correlation structure and Hermite rank of the initialization uniquely determine the critical scaling and asymptotic limit, making them meaningful hyperparameters in the asymptotic regime, whereas finite‑variance i.i.d. initialization always yields a Brownian driver regardless of distribution.

By Felix Benning, Ivan Nourdin, Giovanni Peccati
arXiv Machine Learning
Jul 8

A Functional-Space Mean-Field Theory of Partially-Trained Three-Layer Neural Networks

arXiv:2210. 16286v2 Announce Type: replace Abstract: To understand the training dynamics of neural networks, prior studies have considered the mean-field limit of two-layer neural networks as the width tends to infinity, establishing theoretical guarantees for its convergence under gradient flow training as well as approximation and generalization capabilities.

By Zhengdao Chen, Eric Vanden-Eijnden, Joan Bruna