The paper investigates the relationship between discrete gradient descent (GD) and its continuous-time gradient-flow counterpart in the context of ReLU neural networks. It shows that while GD states converge over a finite horizon, the exact discrete derivatives obtained via automatic differentiation do not necessarily match the derivative of the limiting flow, due to singular curvature at activation events. The authors provide a Stieltjes representation that separates continuous regional Hessians from atomic interface curvature, revealing rank-one discrepancies at activation jumps and demonstrating that even globally strongly convex residual-ReLU losses can exhibit large sensitivity ratios on certain initialization sets.
By Xiaoyang Li, Runni Zhou
The paper investigates how exact automatic differentiation behaves under hard‑ReLU gradient descent. It shows that while gradient‑descent states converge to the piecewise‑smooth gradient flow, the derivative of the training map does not, due to missing event‑time sensitivities captured by saltation matrices. The study demonstrates that for convex objectives, activation events can create large sensitivity gaps, and provides empirical evidence that event‑aware corrections are necessary for accurate flow derivatives.
By Xiaoyang Li, Runni Zhou
The paper introduces a certified continuation framework for computing and training deep equilibrium networks (DEQs). It uses compact input homotopy and a rounded Newton tracker for inference, and augments local-plus-low-rank recurrence with programmable dormant bilinear rank‑one channels for training. The approach guarantees polynomial‑time bit complexity, with certified bounds on inference and training error budgets.
By Alex Borisevich
arXiv:2607. 07665v1 Announce Type: new Abstract: Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules.
By Shiheng Zhang
arXiv:2609. 01319v1 Announce Type: cross Abstract: At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center.
By Ziqi Zhao, Qingjian Ni
Classifier-free guidance (CFG) is the standard way to strengthen class-conditioning in diffusion and flow-matching samplers, yet at large guidance it oversaturates and destabilizes, symptoms practitioners suppress with more steps or limited-interval schedules. We analyze CFG through an asymptotic-preserving, numerical-analysis lens.
arXiv:2607. 20171v1 Announce Type: cross Abstract: Learned solvers for compressible flow are usually compared to classical methods at equal mesh resolution rather than at equal computational cost, and they typically offer no guarantee that their solutions remain physically admissible.
By Denis Gueyffier (ONERA -- Institut Polytechnique de Paris)
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2605. 01928v2 Announce Type: replace Abstract: We optimize losses that jump: spiking thresholds, quantized layers, and discrete routing put jumps in the forward pass, where backpropagation does not apply.
By An T. Le
arXiv:2608. 12655v1 Announce Type: new Abstract: A flat training curve does not reveal whether a neural network has reached a global optimum, is locally trapped, is representation-limited, or is mismatched to its trainer.
By Farhang Yeganegi, Arian Eamaz, Mojtaba Soltanalian
At a junction, a score field can reveal weighted tangent rays, yet these first-order quantities do not determine how individual branches bend or how their densities change away from the center. Recove...
arXiv:2608. 00737v1 Announce Type: new Abstract: Hard-constrained recurrent physics-informed networks (HRPINNs) embed known dynamics inside a recurrent numerical integrator and restrict a neural branch to learning only the residual dynamics that the first-principles model does not capture.
By Enzo Nicolas Spotorno, Josafat Leal Filho