arXiv Machine Learning

Weak Correlations as the Underlying Principle for Linearization of Gradient-Based Learning Systems

arXiv:2401. 04013v2 Announce Type: replace Abstract: Deep learning models, such as wide neural networks, can be conceptualized as nonlinear dynamical physical systems characterized by a multitude of interacting degrees of freedom.

arXiv Machine Learning
Aug 11

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

By Andrea Combette, Nelly Pustelnik, Antoine Venaille
arXiv Machine Learning
Sep 10

The Dynamics of Generalization in Deep Learning

arXiv:2504.16450v4 Announce Type: replace Abstract: We derive a differential equation that governs the evolution of the generalization gap when a model is trained by gradient descent-based methods. T...

By Rubing Yang, Pratik Chaudhari
arXiv Machine Learning
Jul 30

Minimax-Optimal Generalization Bounds for Smooth Deep Neural Networks Trained by (Stochastic) Gradient Descent

arXiv:2606. 06772v2 Announce Type: replace-cross Abstract: Characterizing the optimization dynamics and statistical performance of over-parameterized deep neural networks (DNNs) remains a central challenge in understanding the remarkable success of deep learning.

By Junyu Zhou, Puyu Wang, Dennis Wagner, Yunwen Lei, Marius Kloft, Yiming Ying
arXiv AI
Jun 29

Derivation of effective gradient flow equations and dynamical truncation of training data in Deep Learning

arXiv:2501. 07400v2 Announce Type: replace-cross Abstract: We derive explicit equations governing the cumulative biases and weights in Deep Learning with ReLU activation function, based on gradient descent for the Euclidean loss in the input layer, and under the assumption that the weights are, in a precise sense, adapted to the coordinate system distinguished by the activations.

By Thomas Chen
arXiv Machine Learning
Sep 25

An Analytical Theory of Auxiliary Learning

The paper presents an analytical theory of auxiliary learning, an optimization paradigm where a neural network’s performance on a target task is enhanced by jointly training on additional tasks. Using a teacher‑student framework, the authors derive a closed system of differential equations that describe online stochastic gradient descent dynamics in the large‑input limit. For linear networks, they provide a closed‑form expression for the generalization error that shows how task correlations and label noise influence the benefit of auxiliary learning, while for nonlinear activations they develop a fluctuation‑dissipation theory linking main, auxiliary, and single‑task errors. Numerical experiments confirm the theory and illustrate how auxiliary tasks improve generalization by balancing forcing dynamics toward the optimal solution with gradient noise.

By Federico Milanesio, Alessandro Ingrosso, Matteo Osella
arXiv Machine Learning
Sep 25

Pointwise Generalization in Deep Neural Networks

The paper introduces a pointwise generalization theory for fully connected deep neural networks, using a pointwise Riemannian Dimension derived from eigenvalues of learned feature representations across layers. This framework provides hypothesis-dependent, representation-aware generalization bounds that are significantly tighter than traditional size- or norm-based approaches, both theoretically and experimentally. The authors analytically identify structural properties that explain deep networks’ tractability and empirically show that the pointwise Riemannian Dimension captures feature compression, over‑parameterization effects, and optimizer bias.

By Shaojie Li, Yunbei Xu