arXiv Machine Learning By Andy Arditi, Weian Xie, David Bau, Liu Ziyin

Learning the identity: a case study of how SGD selects among functional decompositions

Read the original on arXiv Machine Learning →

The paper investigates how stochastic gradient descent (SGD) selects specific functional decompositions when training a deep linear residual network to learn the identity function. Although many weight configurations minimize the population loss, SGD consistently prefers particular solutions, especially under anisotropic label noise or different parametrizations. The authors explain this bias using an entropic loss term that penalizes the expected squared norm of the minibatch gradient, analytically characterizing its minimizers and showing that trained networks align with these predictions.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
2d ago

Exact information accounting for SGD methods

arXiv:2610.00446v1 Announce Type: cross Abstract: As an alternative to the standard geometric analyses, we give an exact, information-theoretic analysis of stochastic gradient descent (SGD) and its v...

By Akshay Balsubramani
arXiv Machine Learning
Jun 19

Fisher-Geometric Sharpness and the Implicit Bias of SGD toward Flat Minima

arXiv:2606. 20469v1 Announce Type: new Abstract: A widely held intuition in deep learning is that stochastic gradient descent (SGD) implicitly favors flat minima and that flat minima generalize better, but standard Euclidean measures of flatness such as the trace or maximum eigenvalue of the loss Hessian are not invariant under reparametrizations that preserve the network function, which undermines the theoretical foundations of this narrative.

By Md Sakir Ahmed, Kumaresh Sarmah, Hemen Dutta
arXiv Machine Learning
Jul 7

Learning rate adaptive stochastic gradient descent optimization methods: numerical simulations for deep learning methods for partial differential equations and convergence analyses

arXiv:2406. 14340v2 Announce Type: replace-cross Abstract: The standard stochastic gradient descent (SGD) optimization method, as well as adaptive methods such as the Adam optimizer fail to converge if the learning rates do not converge to zero (particularly, in the situation of constant learning rates).

By Steffen Dereich, Arnulf Jentzen, Adrian Riekert