arXiv:2607. 14576v1 Announce Type: new Abstract: We propose \emph{the sublinear-growth principle} for deep residual architectures -- a sharp stability threshold on the input-magnitude exponent of every residual block's velocity field: $$\|v(x, t)\| \leq c\,\|x\|^q + b, \qquad q \in [0, 1].
By Hyemin Gu, Michael Tyrrell, Tuhin Sahai, Markos A. Katsoulakis
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2602. 18849v2 Announce Type: replace-cross Abstract: We develop a sensitivity analysis for transformer attention in a geometry aligned with tokenwise computation.
By Seyed Morteza Emadi
arXiv:2607. 23390v1 Announce Type: new Abstract: When can additional low-bit residual computation replace missing numerical precision for a fixed input-output map?
By Mojtaba Soltanalian
The paper introduces a certified continuation framework for computing and training deep equilibrium networks (DEQs). It uses compact input homotopy and a rounded Newton tracker for inference, and augments local-plus-low-rank recurrence with programmable dormant bilinear rank‑one channels for training. The approach guarantees polynomial‑time bit complexity, with certified bounds on inference and training error budgets.
By Alex Borisevich
arXiv:2604. 07328v3 Announce Type: replace Abstract: How does the choice of training data influence an AI model?
By Sam Gunn
The paper introduces a new closed‑form scaling law that extends Chinchilla’s original formula to handle data‑constrained regimes. It decomposes loss into undercapacity, undertraining, and overfitting components, saturating between an irreducible loss and an uninformed baseline. The authors validate the model on diverse architectures and domains, achieving state‑of‑the‑art RMSE across multiple LLM scaling‑law grids and enabling cost‑aware training allocations.
By Christopher M. Bryant, Hao Liu
arXiv:2605. 25085v2 Announce Type: replace-cross Abstract: We study the rate-distortion limits of online KV cache compression in autoregressive language models, formulating it as sequential Wyner-Ziv source coding on the filtration induced by the model, with the next-step query as decoder side information.
By Munsik Kim
arXiv:2609.39813v1 Announce Type: new
Abstract: Low-precision training rounds tensors that the backward pass reads again, often for several gradients; each use can read the forward's rounded value, t...
By Shuxiao Xie, Shuyang Xie, Dezhi Ran, Wei Yang, Tao Xie
arXiv:2606. 04058v1 Announce Type: cross Abstract: Orthonormalized update rules have rapidly become a leading choice of optimizer for training large language models, with recent open-source state-of-the-art models adopting Muon.
By Gagik Magakyan, Pablo Parrilo, Asuman Ozdaglar
arXiv:2605.06240v2 Announce Type: replace-cross
Abstract: Forward-Forward (FF) training lets each layer learn from a local goodness criterion. In cumulative-goodness variants, later layers can inheri...
By Amirhossein Yousefiramandi
arXiv:2605. 24033v2 Announce Type: replace Abstract: Mechanistic interpretability typically discovers circuits and then argues what they do from examples and ablations.
By Neel Somani