SpanNorm: Reconciling Training Stability and Performance in Deep Transformers
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.
arXiv:2601. 22580v2 Announce Type: replace-cross Abstract: The success of Large Language Models (LLMs) hinges on the stable training of deep Transformer architectures.
arXiv:2607. 10593v1 Announce Type: new Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent.
The paper introduces Anon, an optimizer that extends adaptivity beyond the traditional bounds of SGD and Adam by allowing extrapolation across the entire real-number spectrum. It addresses the limitations of prior tunable optimizers that only interpolate between 0 and 1 adaptivity, showing that optimal adaptivity can require negative values for CNNs or values greater than one for Transformers. Anon incorporates Incremental Delay Update (IDU) to maintain provable stability and demonstrates competitive performance on image classification, diffusion, and large language modeling tasks.
arXiv:2608. 16760v1 Announce Type: new Abstract: Reliable optimization is central to neural network (NN) training, yet Adam, the default optimizer for modern LLMs, rests on a fragile foundation.
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.
arXiv:2606. 14187v1 Announce Type: new Abstract: Large-scale neural network training increasingly relies on matrix-aware optimizers that exploit the structure of weight parameters beyond element-wise adaptation.
arXiv:2609.15975v1 Announce Type: cross Abstract: Transformer representations evolve through learned additive transformations that either preserve their current direction or redirect it. We study thi...
arXiv:2607. 09967v1 Announce Type: cross Abstract: Many neural networks operations have a multiplicative nature rather than additive: halving or doubling a norm are analogous relatively but require unequal optimization distances when taking linear steps.
arXiv:2601. 17257v2 Announce Type: replace Abstract: We introduce a constrained optimization framework for training transformers that behave like optimization descent algorithms.
arXiv:2608.30720v1 Announce Type: new Abstract: Representational similarity is foundational to analyses of deep networks, yet distances between point-valued representations are not intrinsically tied...
arXiv:2606. 26538v1 Announce Type: cross Abstract: Deep Transformers are composed of uniformly stacked residual blocks, yet their deepest layers often add little value.