arXiv AI

A Link between Shock-wave Theory and Symmetry-reduced Stochastic Gradient Descent for Artificial Neural Networks

arXiv:2606. 18303v1 Announce Type: cross Abstract: We develop a mathematically explicit link between shock-wave theory and the symmetry-quotiented learning dynamics of stochastic gradient descent, drawing on differential geometry, Lie group theory, and fluid mechanics.

arXiv AI
Aug 6

The Hamilton-Jacobi Theory of Deep Learning

arXiv:2605. 28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights.

By Jose Marie Antonio Mi\~noza, Erika Fille T. Legara, Christopher P. Monterola
arXiv Machine Learning
Sep 3

Emergence of Fibrations, Compression, and Symmetry Breaking in Artificial Neural Networks

Artificial neural networks generate local symmetries called fibrations and coverings during learning, and these covering symmetries are stable attractors of stochastic gradient descent. The study shows that such symmetries appear across diverse architectures—multilayer, convolutional, recurrent, and transformer networks—and can be exploited for drastic model compression, reducing networks to 17% of their original size without performance loss. Controlled breaking of covering symmetry further improves continual learning, achieving state‑of‑the‑art results.

By Osvaldo M Velarde, Lucas C Parra, Alireza Hashemi, Hernan A Makse
arXiv Machine Learning
Sep 3

Learning and extrapolating scale-invariant processes

The paper investigates how machine learning models can regress scale‑free processes, such as earthquakes or avalanches, focusing on predicting rare, large events that require extrapolation. It studies two self‑similar systems: a 2‑dimensional fractional Gaussian field and the Abelian sandpile model. Experiments compare existing architectures (U‑net, Riesz network) with new proposals (wavelet‑based Graph Neural Network, Fourier embedding, Fourier‑Mellin Neural Operator) to identify spectral bias and coarse‑graining challenges and suggest inductive biases to address them.

By Anaclara Alvez-Canepa, Cyril Furtlehner, Fran\c{c}ois P. Landes