arXiv:2511. 07308v3 Announce Type: replace Abstract: Understanding the training dynamics of deep neural networks remains a major open problem, with physics-inspired approaches offering promising insights.
By Ildus Sadrtdinov, Ekaterina Lobacheva, Ivan Klimov, Mikhail Burtsev, Mikhail I. Katsnelson, Dmitry Vetrov
arXiv:2608. 01833v1 Announce Type: cross Abstract: Grokking is a striking phenomenon in neural network training, where a model can undergo a prolonged period of pure memorization before abrupt generalization.
By Lai Shun Chan, Xiaotian Zhang, Yue Shang, Ge Zhang, Entao Yang
arXiv:2606. 04476v1 Announce Type: new Abstract: In this paper, we study the gradient descent dynamics for jointly training both layers of a one-hidden-layer ReLU network to fit a linear target function.
By Berk Tinaz, Changzhi Xie, Mahdi Soltanolkotabi
arXiv:2607. 29503v1 Announce Type: new Abstract: While neural networks are typically evaluated by their training and test performance, these metrics do not reveal how robust a learned representation is.
By Xiaotian Zhang, Lai Shun Chan, Yue Shang, Entao Yang, Ge Zhang
arXiv:2503. 07325v2 Announce Type: replace Abstract: Understanding and certifying the behavior of modern deep neural networks remains a fundamental challenge in reliable machine learning.
By Khoat Than, Dat Phan
arXiv:2606. 20299v1 Announce Type: cross Abstract: Deep learning has managed to evade numerous intuitions from classical statistics to achieve unprecedented performance on a number of real-world tasks.
By Itay Lavie, Noam Levi, Yonatan Kahn