arXiv Machine Learning By Lorenzo Livi

Anti-Collapse Dynamics and the Emergence of Multi-Time-Scale Learning in Recurrent Neural Networks

Read the original on arXiv Machine Learning →

arXiv:2606. 29519v1 Announce Type: new Abstract: Long-range learning is hard for recurrent networks trained with stochastic gradient descent, because the influence of a past input fades with the lag $\ell$, and if it fades too fast the dependence cannot be learned from finite data.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.