The paper introduces a finite‑width geometric framework that explains how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. It quantifies incompatibilities among weight‑generated covariance, gates, and backward sensitivities using three families of commutators, and provides exact layerwise identities that decompose these commutators into sources such as downstream transport, adjacent‑layer imbalance, and nonlinear gate‑covariance interactions. The study demonstrates that spectral alignment is a layer‑ and scale‑dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.
By Kaj Nystr\"om
The paper investigates how exact automatic differentiation behaves under hard‑ReLU gradient descent. It shows that while gradient‑descent states converge to the piecewise‑smooth gradient flow, the derivative of the training map does not, due to missing event‑time sensitivities captured by saltation matrices. The study demonstrates that for convex objectives, activation events can create large sensitivity gaps, and provides empirical evidence that event‑aware corrections are necessary for accurate flow derivatives.
By Xiaoyang Li, Runni Zhou
arXiv:2603. 13751v2 Announce Type: replace Abstract: Physics-informed neural networks (PINNs) have achieved notable success in modeling dynamical systems governed by partial differential equations (PDEs).
By Zhangyong Liang, Huanhuan Gao
arXiv:2607. 17696v1 Announce Type: cross Abstract: We develop an adjoint-sensitivity framework for positional influence in causal residual Transformers and separate unconditional analytic results from conditional boundary-shape conclusions.
By Cheng Huan, Hongwei Yuan
The paper introduces COFM, a framework for consistent optimal transport flow matching that uses partially input convex neural networks (PICNN) to parameterize the transport potential. By adding a Hamilton‑Jacobi residual to the training objective, COFM enforces dynamical consistency and supports both one‑step transport and multi‑step ODE sampling without costly inner optimization. Experiments on benchmark datasets show that COFM achieves competitive performance while reducing L^2‑UVP by over 2× and cutting computational time by about 9× compared to state‑of‑the‑art models.
By Fanghui Song, Zhongjian Wang, Jiebao Sun
arXiv:2606. 06171v1 Announce Type: cross Abstract: Physics-Informed Neural Networks inherently suffer from task interference because they rely on a shared parameter space to satisfy both governing differential equations and boundary conditions.
By Cornelius Otchere, Michael Shields
The paper investigates how gradient descent behaves near codimension‑one bifurcations in recurrent neural networks by analyzing the global empirical Neural Tangent Kernel (GeNTK). Under local center‑manifold conditions, the parameter‑to‑state Jacobian is approximated by a low‑rank normal‑form operator, causing the GeNTK and Fisher information matrix to become strongly amplified and anisotropic, concentrating on a rank‑one or rank‑two channel depending on the bifurcation type. Experiments on high‑dimensional RNNs confirm that this low‑rank concentration coincides with abrupt loss changes, subtask interference, and aligns with changes in memory dynamics in a 15‑task LeakyRNN.
By James Hazelden, Eric Shea-Brown
We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization. We reinterpret skip connections and normalization, long understood as controlling magnitude, as mechanisms for preserving gradient rank across depth, since the very matrix multiplications and nonlinear activations that make the network expressive also reduce the rank.
arXiv:2609.08381v1 Announce Type: cross
Abstract: Equivariant networks are commonly trained with Adam, yet recent work reports that matrix-structured optimizers such as Muon can perform better on the...
By Andrei Manolache, Mathias Niepert
arXiv:2606. 15892v1 Announce Type: new Abstract: Accurate interatomic potentials enable molecular dynamics of materials, molecules, and interfaces beyond density-functional-theory length and time scales.
By Jia Bi, Alin Marin Elena, Samuel Pinilla
arXiv:2605. 28983v2 Announce Type: replace-cross Abstract: In this paper, training a neural network is identified, exactly, as a search through Hamilton--Jacobi initial-value problems: each gradient step selects the initial data of a viscous Hamilton--Jacobi equation whose Hopf--Cole propagator best fits the observations; at inference, the input is the spatial point at which that solution is evaluated and the initial condition is already encoded in the weights.
By Jose Marie Antonio Mi\~noza, Erika Fille T. Legara, Christopher P. Monterola
arXiv:2607. 14018v1 Announce Type: cross Abstract: We investigate how each component of the Transformer feedforward block architecture design determines how much rank survives across depth at initialization.
By Katie Everett