arXiv:2609.16827v1 Announce Type: new
Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit exceptional storage capabilities and robustness. Previous empirica...
By Akira Tamamori
arXiv:2605. 00366v4 Announce Type: replace-cross Abstract: High-capacity associative memories based on Kernel Logistic Regression (KLR) exhibit strong storage capabilities, but the dynamical and geometric mechanisms underlying their stability remain poorly understood.
By Akira Tamamori
arXiv:2511. 01938v3 Announce Type: replace-cross Abstract: Grokking is a puzzling phenomenon in neural networks where full generalization occurs only after a substantial delay following the complete memorization of the training data.
By Tiberiu Musat
arXiv:2609.35948v1 Announce Type: cross
Abstract: Geometry does more than constrain an associative memory: curvature determines what it remembers and which states it creates. We develop intrinsic den...
By Krishnakumar Balasubramanian, Zhaoyang Shi
arXiv:2606. 04212v1 Announce Type: new Abstract: Existing analyses of the edge of stability (EoS) treat it as a global property of optimization.
By Shauna Kwag, Anakha Ganesh, Tomaso Poggio, Pierfrancesco Beneventano
arXiv:2607. 00947v1 Announce Type: new Abstract: Generative models learn data distributions that reside on a low-dimensional manifold within a higher-dimensional ambient space.
By Ludwig Winkler, Andrew Leaver-Fay, Joseph Kleinhenz, Pan Kessel
arXiv:2606. 30226v1 Announce Type: new Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics.
By Marcelina Marjankowska, Valerio Modugno, Paolo Barucca
arXiv:2606. 30388v1 Announce Type: cross Abstract: Delayed generalization (\ie~grokking) refers to the phenomenon in which a neural network fits its training data early in training but only begins to generalize after a prolonged delay, often through an abrupt transition.
By R\'ois\'in Luo, Christian Gagn\'e, Jonas Ngnaw\'e, Ihsan Ullah, Karyn Morrissey
arXiv:2608. 04382v1 Announce Type: new Abstract: Gradient descent has been of particular interest in modern machine learning beyond sole focus on optimization.
By Han Bao
arXiv:2602. 14789v2 Announce Type: replace Abstract: The dynamical stability of the iterates during training plays a key role in determining the minima obtained by optimization algorithms.
By Rotem Mulayoff, Sebastian U. Stich
arXiv:2606. 30512v1 Announce Type: cross Abstract: Why overparameterised deep networks generalise so remarkably well remains one of the most stubborn open questions in machine learning theory.
By Srinivasa Rao P., Vangmayi P Reddy
arXiv:2606. 02993v1 Announce Type: new Abstract: Understanding how structured internal structure emerges during neural network training is central to the study of deep learning.
By Jianliang He, Leda Wang, Fengzhuo Zhang, Siyu Chen, Zhuoran Yang