arXiv:2606. 17816v1 Announce Type: cross Abstract: Understanding gradient descent dynamics is key to explaining the success of over-parameterized models, where implicit bias manifests through conservation laws in gradient flow.
By Viet-Hoang Tran, Vinh Khanh Bui, Tan Lai Ngoc, Nam Nguyen, Tuan Dam, Tan M. Nguyen
arXiv:2606. 04754v1 Announce Type: new Abstract: Many striking phenomena in deep learning, such as linear mode connectivity and the structured behavior of training dynamics, are closely tied to parameter symmetries: transformations that leave the realized function unchanged.
By Vincent B\"urgin, Daniel Herbst, Ya-Wei Eileen Lin, Stefanie Jegelka
arXiv:2609.39078v1 Announce Type: new
Abstract: Representations are routinely used across machine learning, psychology, and neuroscience to draw inferences about the computations of biological and ar...
By Marvin Theiss, Lukas Braun, Andrew M. Saxe, Erin Grant
arXiv:2604. 14037v2 Announce Type: replace Abstract: Parameter space is not function space for neural network architectures.
By Pranavkrishnan Ramakrishnan
Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields. Despite their remarkable capacity for representing geometric structures, ENNs suffer from degraded expressivity when processing symmetric inputs: the output representations are invariant to transformations that extend beyond the input's symmetries.
arXiv:2608. 12010v1 Announce Type: new Abstract: Equivariant Neural Networks (ENNs) have empowered numerous applications in scientific fields.
By Ning Lin, Jiacheng Cen, Anyi Li, Wenbing Huang, Hao Sun
Artificial neural networks generate local symmetries called fibrations and coverings during learning, and these covering symmetries are stable attractors of stochastic gradient descent. The study shows that such symmetries appear across diverse architectures—multilayer, convolutional, recurrent, and transformer networks—and can be exploited for drastic model compression, reducing networks to 17% of their original size without performance loss. Controlled breaking of covering symmetry further improves continual learning, achieving state‑of‑the‑art results.
By Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi, Hernan A Makse
arXiv:2607. 07845v1 Announce Type: new Abstract: The Hessian of the training loss governs the local geometry of the loss landscape, yet despite existing explanations for its largest eigenvalues, the origin of the vast multitude of vanishingly small eigenvalues remains elusive.
By Marcel K\"uhn, Bernd Rosenow
arXiv:2606. 18303v1 Announce Type: cross Abstract: We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics.
By Taiki Miyagawa
The paper introduces a diagnostic tool that measures how neural emulators of partial differential equations capture physical symmetries by evaluating the overlap of loss gradients along symmetry-related states. This metric probes the local geometry of the learned loss landscape and goes beyond traditional equivariance tests by directly assessing learning dynamics. Applied to autoregressive fluid flow emulators, the study shows that orbit-wise gradient coherence enables generalization over symmetry transformations and reveals when training selects a symmetry-compatible basin.
By James Amarel, Robyn Miller, Nicolas Hengartner, Benjamin Migliori, Emily Casleton, Alexei Skurikhin, Earl Lawrence, Gerd J. Kunde
arXiv:2501. 07400v2 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations.
By Thomas Chen
CrossGMN introduces a graph metanetwork that processes a trained source network and an initialized target network simultaneously, enabling equivariant cross‑architecture weight‑space transformations. By preserving symmetry through cross‑network message passing, CrossGMN can refine target network initializations while remaining invariant to source permutations and equivariant to target permutations. Experiments demonstrate that CrossGMN accelerates knowledge distillation, transfers across datasets without retraining, and unifies compression from diverse source architectures into a common target architecture.
By Adir Dayan, Yam Eitan, Haggai Maron