arXiv Machine Learning

Hard-ReLU Gradient Descent Selects an Event-Free Sensitivity Limit

The paper investigates how exact automatic differentiation behaves under hard‑ReLU gradient descent. It shows that while gradient‑descent states converge to the piecewise‑smooth gradient flow, the derivative of the training map does not, due to missing event‑time sensitivities captured by saltation matrices. The study demonstrates that for convex objectives, activation events can create large sensitivity gaps, and provides empirical evidence that event‑aware corrections are necessary for accurate flow derivatives.

arXiv Machine Learning
Sep 1

Singular Curvature in ReLU Training:Differentiation and the Gradient-Flow Limit Need Not Commute

The paper investigates the relationship between discrete gradient descent (GD) and its continuous-time gradient-flow counterpart in the context of ReLU neural networks. It shows that while GD states converge over a finite horizon, the exact discrete derivatives obtained via automatic differentiation do not necessarily match the derivative of the limiting flow, due to singular curvature at activation events. The authors provide a Stieltjes representation that separates continuous regional Hessians from atomic interface curvature, revealing rank-one discrepancies at activation jumps and demonstrating that even globally strongly convex residual-ReLU losses can exhibit large sensitivity ratios on certain initialization sets.

By Xiaoyang Li, Runni Zhou
arXiv Statistics ML
Aug 25

A Commutator Framework for Selective Spectral Alignment in Deep Neural Networks

The paper introduces a finite‑width geometric framework that explains how learned feature geometries are organized, transported, and selectively aligned in deep neural networks. It quantifies incompatibilities among weight‑generated covariance, gates, and backward sensitivities using three families of commutators, and provides exact layerwise identities that decompose these commutators into sources such as downstream transport, adjacent‑layer imbalance, and nonlinear gate‑covariance interactions. The study demonstrates that spectral alignment is a layer‑ and scale‑dependent compatibility phenomenon governed by transport, interaction, cancellation, and possible damping, rather than a universal consequence of training.

By Kaj Nystr\"om
arXiv Machine Learning
Sep 10

Branch Geometry and Finite-Radius Sensitivity of Hard-ReLU Training

The paper investigates how hard‑ReLU training behaves when perturbations have a finite radius. It shows that the usual infinitesimal sensitivities are insufficient to predict the response at a chosen radius, and it characterizes the intermediate regime where the perturbation radius scales with the gradient‑descent step. The authors derive crossing indices, a uniform endpoint expansion for separated transverse events, and provide explicit remainder terms in contractive affine regions to certify finite candidate comparisons, supported by experiments on nonlinear networks.

By Xiaoyang Li, Runni Zhou, Xinghao Yan
arXiv Machine Learning
Jun 2

Multigrade Neural Network Approximation

arXiv:2601. 16884v3 Announce Type: replace Abstract: We study multigrade deep learning (MGDL) as a principled framework for structured error refinement in deep neural networks.

By Shijun Zhang, Zuowei Shen, Yuesheng Xu