arXiv:2602. 05779v2 Announce Type: replace Abstract: The Edge-of-Chaos (EoC) theory developed for the random initialization of deep networks allows more efficient training by both preserving information in the initial outputs of the network and minimising exploding or vanishing gradients through characterisation of the intermediate layers as Gaussian processes.
By Emily Dent, Jared Tanner
arXiv:2506. 08764v3 Announce Type: replace Abstract: Deep neural networks are known to suffer from exploding or vanishing gradients as depth increases, a phenomenon closely tied to the spectral behavior of the input-output Jacobian.
By Benjamin Dadoun, Soufiane Hayou, Hanan Salam, Mohamed El Amine Seddik, Pierre Youssef
arXiv:2507.10383v5 Announce Type: replace-cross
Abstract: Recurrent neural networks are canonical models of biological memory. In these models, memories are represented by distributed patterns of neu...
By Uri Cohen, M\'at\'e Lengyel
The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matching and derivative regularization depend on higher...
arXiv:2609.09244v1 Announce Type: new
Abstract: The edge-of-chaos condition preserves first-order input perturbations in wide randomly initialized networks, but physics-informed losses, score matchin...
By Prashant Singh, Pranav Singh
arXiv:2512.12767v2 Announce Type: replace-cross
Abstract: Training recurrent neuronal networks consisting of excitatory (E) and inhibitory (I) units with additive noise for working memory computation...
By Thiparat Chotibut, Oleg Evnin, Weerawit Horinouchi
arXiv:2606. 07120v1 Announce Type: new Abstract: Autoencoders (AEs) learn low-dimensional representations by mapping data into a latent space while minimizing reconstruction error.
By Santanu Das, Ramyak Bilas, Pascal Esser, Satyaki Mukherjee
arXiv:2606. 05326v1 Announce Type: cross Abstract: We study the dynamics of gradient descent in the Edge of Stability regime, where the learning rate is large enough to induce persistent oscillations in the loss and the sharpness.
By Antonin Chodron de Courcel
arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.
By Andrea Combette, Nelly Pustelnik, Antoine Venaille
arXiv:2602. 10949v2 Announce Type: replace-cross Abstract: Effective initialization in deep networks requires an understanding of random neural networks.
By Constantin Kogler, Tassilo Schwarz, Samuel Kittle
arXiv:2609.01034v1 Announce Type: new
Abstract: The central flow of Cohen et al. (2025) is an empirically accurate continuous-time model of gradient descent at the edge of stability in deep learning,...
By Rapha\"el Berthier
The paper investigates how parameter Jacobians influence the stability of network outputs within the framework of network dynamics, learning models, and neural tangent kernels (NTK). It demonstrates that linearized dynamics can be expressed as a semigroup of linear operators on Hilbert spaces, and provides explicit a priori perturbation bounds for fixed‑kernel linearizations in the NTK setting. The authors also offer refinements for task‑specific spaces, ergodic comparison estimates, spectral‑distribution conditions, and extensions to nonautonomous NTK evolutions, supported by worked examples.
By Halyun Jeong, Palle E. T. Jorgensen, Hyun-Kyoung Kwon, Myung-Sin Song, James Tian