arXiv Machine Learning By Parsa Rangriz

Limit Theorems for Stochastic Gradient Descent in High-Dimensional Single-Layer Networks

Read the original on arXiv Machine Learning →

arXiv:2511. 02258v3 Announce Type: replace-cross Abstract: This paper studies the high-dimensional scaling limits of online stochastic gradient descent (SGD).

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

Hugging Face Trending Papers
Jul 5

Broken Ergodicity and the Violation of the Fluctuation-Dissipation Theorem Lead to Generalization Beyond Overfitting in Machine Learning

The remarkable ability of modern neural networks to generalize improves with increasing network capacity, even when the number of model parameters or effective degrees of freedom exceeds the number of training data points. This phenomenon is all the more surprising given that generalization error diverges when the number of model parameters approaches a critical value from below.