arXiv Machine Learning By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang

Tunneling the Loss Landscape: Bypassing Memorization with Monte Carlo Parameter Swapping

Read the original on arXiv Machine Learning →

arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Jul 1

Revisiting the Volume Hypothesis

arXiv:2606. 31282v1 Announce Type: new Abstract: Modern deep neural networks often contain far more parameters than needed to fit their training data, yet they achieve impressive generalization.

By Ari Pakman, Lior Kreimer, Yakir Berchenko
arXiv AI
Jun 30

A Stochastic--Geometric Theory of Scaling Laws in Grokking

arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.

By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey