arXiv Machine Learning By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

Read the original on arXiv Machine Learning →

arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.