arXiv:2609.39078v1 Announce Type: new
Abstract: Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and ar...
By Marvin Theiss, Lukas Braun, Andrew M. Saxe, Erin Grant
arXiv:2604. 14037v2 Announce Type: replace Abstract: Parameter space is not function space for neural network architectures.
By Pranavkrishnan Ramakrishnan
arXiv:2606. 10913v1 Announce Type: new Abstract: We explore whether intrinsic symmetries of the training data lead to conserved quantities during gradient-flow training of neural networks.
By Jakob Galley, Vahid Shahverdi, Axel Flinth
Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.
arXiv:2608. 12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
By Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
Artificial neural networks generate local symmetries called fibrations and coverings during learning, and these covering symmetries are stable attractors of stochastic gradient descent. The study shows that such symmetries appear across diverse architectures—multilayer, convolutional, recurrent, and transformer networks—and can be exploited for drastic model compression, reducing networks to 17% of their original size without performance loss. Controlled breaking of covering symmetry further improves continual learning, achieving state‑of‑the‑art results.
By Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi, Hernan A Makse
arXiv:2608.24700v1 Announce Type: new
Abstract: When a network has learned a function with a known symmetry, can that symmetry be moved through the parametrisation---is there a motion in parameter sp...
By Alan Muriithi, Vedanta Thapar, Torben Berndt
arXiv:2607. 07845v1 Announce Type: new Abstract: The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive.
By Marcel K\"uhn, Bernd Rosenow
arXiv:2511. 09432v2 Announce Type: replace Abstract: Machine learning (ML) models achieve remarkable performance but remain hard to interpret due to their scale and complexity.
By Ege Erdogan, Ana Lucic
arXiv:2609.38998v1 Announce Type: new
Abstract: A foundational principle of connectionism is that perception, action, and cognition emerge from parallel computations among simple, interconnected unit...
By Lukas Braun, Erin Grant, Andrew M. Saxe
arXiv:2606. 18676v1 Announce Type: new Abstract: Training-free neural architecture search promises efficient discovery of high-performance networks without costly training.
By Qinqin Zhou, Fuhai Chen, Jipeng Wu, Zhiwei Chen, Zhikai Hu, Weiwei Cai
arXiv:2607. 05546v1 Announce Type: cross Abstract: We develop a unified function space theory of deep fully connected neural networks.
By Julia Nakhleh, Robert D. Nowak