arXiv AI

Stability of Transformers under Layer Normalization

arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.

arXiv Machine Learning
Jul 14

AutoNorm: Understanding Adaptive Normalization in Transformers through Differentiable Gating

arXiv:2607. 10593v1 Announce Type: new Abstract: Normalization is a critical component for stabilizing Transformer training, yet the choice between static strategies such as Layer Normalization (LN) and adaptive alternatives remains largely task-dependent.

By Piyush Kaushik Bhattacharyya, Divyanshu Rai, Swastik Singh, Kumar Aakash, Ayush Ranjan, Krutika Verma
arXiv AI
Aug 19

Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization

The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.

By Akiyoshi Tomihari, Issei Sato
arXiv Machine Learning
Sep 2

Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective

The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.

By Ruoxi Yu, Haotian Jiang, Jingpu Cheng, Penghao Yu, Qianxiao Li, Zhong Li