Hugging Face Trending Papers

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

arXiv Statistics ML
Aug 25

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

The paper introduces a finite‑width geometric framework that explains how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. It quantifies incompatibilities among weight‑generated covariance, gates, and backward sensitivities using three families of commutators, and provides exact layerwise identities that decompose these commutators into sources such as downstream transport, adjacent‑layer imbalance, and nonlinear gate‑covariance interactions. The study demonstrates that spectral alignment is a layer‑ and scale‑dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.

By Kaj Nystr\"om
arXiv Machine Learning
Sep 4

Hard-ReLU Gradient Descent Selects an Event-Free Sensitivity Limit

The paper investigates how exact automatic differentiation behaves under hard‑ReLU gradient descent. It shows that while gradient‑descent states converge to the piecewise‑smooth gradient flow, the derivative of the training map does not, due to missing event‑time sensitivities captured by saltation matrices. The study demonstrates that for convex objectives, activation events can create large sensitivity gaps, and provides empirical evidence that event‑aware corrections are necessary for accurate flow derivatives.

By Xiaoyang Li, Runni Zhou
arXiv Machine Learning
Aug 28

COFM: Consistent Optimal Transport Flow Matching via Partially Input Convex Neural Networks

The paper introduces COFM, a framework for consistent optimal transport flow matching that uses partially input convex neural networks (PICNN) to parameterize the transport potential. By adding a Hamilton‑Jacobi residual to the training objective, COFM enforces dynamical consistency and supports both one‑step transport and multi‑step ODE sampling without costly inner optimization. Experiments on benchmark datasets show that COFM achieves competitive performance while reducing L^2‑UVP by over 2× and cutting computational time by about 9× compared to state‑of‑the‑art models.

By Fanghui Song, Zhongjian Wang, Jiebao Sun
arXiv Machine Learning
Aug 19

Center-Manifold Reduction of Learning at Bifurcations: Interference and Rich Learning in Recurrent Neural Networks

The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.

By James Hazelden, Eric Shea-Brown
Hugging Face Trending Papers
Jul 15

Transforming Rank: How Architecture Navigates the Spectral Pathologies of Depth

We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.

arXiv AI
Sep 10

Equivariance Breaks the Learning Rate

arXiv:2609.08381v1 Announce Type: cross Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...

By Andrei Manolache, Mathias Niepert
arXiv AI
Aug 6

The Hamilton-Jacobi Theory of Deep Learning

arXiv:2605. 28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights.

By Jose Marie Antonio Mi\~noza, Erika Fille T. Legara, Christopher P. Monterola