Certifying Residual Architectures from Their Primitives: A Sharp Stability Threshold
Read the original on arXiv Statistics ML →The Flow has not summarised this story yet — read it at arXiv Statistics ML.
The Flow has not summarised this story yet — read it at arXiv Statistics ML.
arXiv:2607. 14576v1 Announce Type: new Abstract: We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: $$\|v(x, t)\| \leq c\,\|x\|^q + b, \qquad q \in [0, 1].
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
arXiv:2607. 23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map?
The paper introduces a certified continuation framework for computing and training deep equilibrium networks (DEQs). It uses compact input homotopy and a rounded Newton tracker for inference, and augments local-plus-low-rank recurrence with programmable dormant bilinear rank‑one channels for training. The approach guarantees polynomial‑time bit complexity, with certified bounds on inference and training error budgets.
arXiv:2604. 07328v3 Announce Type: replace Abstract: How does the choice of training data influence an AI model?