arXiv:2608. 13335v1 Announce Type: new Abstract: Neural networks trained by gradient descent on a smooth cost function can nevertheless learn in steps: the cost holds on long plateaus and then drops abruptly.
By Liu Ziyin, Yizhou Xu, Tomaso Poggio, Isaac Chuang
arXiv:2604. 14037v2 Announce Type: replace Abstract: Parameter space is not function space for neural network architectures.
By Pranavkrishnan Ramakrishnan
arXiv:2605. 01288v3 Announce Type: replace Abstract: In deep networks with small initialization, training exhibits long plateaus separated by sharp feature-acquisition transitions.
By Divit Rawal, Michael R. DeWeese
arXiv:2606. 04754v1 Announce Type: new Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged.
By Vincent B\"urgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka
arXiv:2607. 16720v1 Announce Type: new Abstract: Understanding deep neural networks remains a central challenge in machine learning.
By Haruka Eshima, Makoto Yamada
arXiv:2607. 03613v1 Announce Type: new Abstract: We study the implicit bias of noisy stochastic gradient descent in training wide two-layer ReLU networks for multivariate regression.
By Shuang Liang, Tom Jacobs, Guido Mont\'ufar
arXiv:2606. 30226v1 Announce Type: new Abstract: Hessian spectral properties are a standard tool in analysing neural-network training, with eigenvalues linked to sharpness, generalization, and optimization dynamics.
By Marcelina Marjankowska, Valerio Modugno, Paolo Barucca
arXiv:2606. 10913v1 Announce Type: new Abstract: We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks.
By Jakob Galley, Vahid Shahverdi, Axel Flinth
arXiv:2608. 06597v1 Announce Type: cross Abstract: A scientific theory of deep learning, comprising learning dynamics and statistical properties of learned models, is rapidly gaining attention.
By Bj\"orn Ladewig, Ibrahim Talha Ersoy, Karoline Wiesner
arXiv:2608. 12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
By Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
arXiv:2606. 18303v1 Announce Type: cross Abstract: We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics.
By Taiki Miyagawa
arXiv:2505. 22578v2 Announce Type: replace Abstract: The optimization of neural networks under weight decay remains poorly understood from a theoretical standpoint.
By Etienne Boursier, Matthew Bowditch, Matthias Englert, Ranko Lazic