arXiv Machine Learning

Optimal Initialization in Depth: Lyapunov Initialization and Limit Theorems for Deep Leaky ReLU Networks

arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.

arXiv Machine Learning
Jun 16

How Controlling the Variance can Improve Training Stability of Sparsely Activated DNNs and CNNs

arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.

By Emily Dent, Jared Tanner
arXiv Machine Learning
Aug 11

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

By Andrea Combette, Nelly Pustelnik, Antoine Venaille
arXiv Machine Learning
5d ago

Stable initialization without the CLT

The paper introduces a new method called uniform‑phase initialization for deep neural networks with sine activations, eliminating the need for the Central Limit Theorem and fully decoupling layers. This approach avoids distributional approximation errors and coupling between layers, leading to stable weight initialization. Experiments show that models using this initialization outperform state‑of‑the‑art methods on image and audio fitting tasks and remain competitive without tuning, while also supporting μP width scaling.

By Simon Kuang, Kyle Chickering, Xinfan Lin
arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying