arXiv Machine Learning By Vivek S Borkar

Exploding and vanishing gradients in deep neural networks: the effect of residual connections

Read the original on arXiv Machine Learning →

arXiv:2606. 17013v1 Announce Type: cross Abstract: The well known phenomenon of exploding and vanishing gradients in deep neural networks is analyzed using multiplicative ergodic theory.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 11

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

By Andrea Combette, Nelly Pustelnik, Antoine Venaille