arXiv Machine Learning

Favorability of Loss Landscape with Weight Decay Requires Both Large Overparametrization and Initialization

arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.