Hugging Face Trending Papers

Gradient Flow Dynamics and Implicit Bias of Diagonal Linear Networks under Infinitesimal Initialization

We study the gradient flow dynamics of diagonal linear networks for regression tasks under infinitesimal initialization. Extending Theorem 1 from Pesme & Flammarion (2023), we generalize the analysis to both deep diagonal linear networks and a broader class of two-layer diagonal linear networks (as defined in Definition 4.

arXiv Machine Learning
Aug 11

Correlation flow governs learning at criticality

arXiv:2608. 08350v1 Announce Type: new Abstract: The initialisation of deep neural networks determines whether information and gradients can propagate across depth, yet a unified theory connecting these properties to learning dynamics remains elusive.

By Andrea Combette, Nelly Pustelnik, Antoine Venaille
arXiv Machine Learning
Sep 11

Learning Orthogonal Multi-Index Models Beyond Small Initialization: Incremental Learning, Competitive Dynamics and Symmetry

The paper investigates how two‑layer polynomial‑width neural networks learn orthogonal multi‑index targets under standard initialization. It shows that incremental learning still occurs: the loss decreases sequentially following the Hermite expansion, with lower‑order components learned first. The dynamics also exhibit a competitive reallocation of parameter mass, shifting into the target subspace and concentrating on aligned neurons. The analysis uses a symmetry‑based finite‑width approximation and demonstrates that vanilla gradient descent displays the same qualitative behavior.

By Mo Zhou, Weihang Xu, Simon S. Du, Maryam Fazel
Hugging Face Trending Papers
Jul 13

Backpropagation as a Nilpotent Linear System

Backpropagation is the computational engine of deep learning, yet its mathematical structure is typically treated as a procedural traversal of computational graphs. We present a global operator theory of the \emph{F-adjoint} framework, which reformulates the layerwise backward recursion of an $L$-depth feedforward network into a single linear system $(I-\cB)\Xs=\bG$, where $\bG$ is a source vector.