Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Related stories
Hugging Face and Graphcore partner for IPU-optimized Transformers
LiFT: Local Search via Linear Programming for Overfitting-Controlled Transformers
arXiv:2606. 16243v1 Announce Type: new Abstract: This paper proposes a Linear Programming (LP)-based local search framework for fine-tuning pretrained transformer models with explicit control against overfitting.
Stability of Transformers under Layer Normalization
arXiv:2510. 09904v2 Announce Type: replace-cross Abstract: Despite their widespread use, training deep Transformers can be unstable.
A Gentle Introduction to 8-bit Matrix Multiplication for transformers at scale using transformers, accelerate and bitsandbytes
Mixture of Experts (MoEs) in Transformers
Convert Transformers to ONNX with Hugging Face Optimum
On the Expressive Power of Transformers
arXiv:2608. 12671v1 Announce Type: new Abstract: Multi-layer transformers form the critical component of essentially all large language models (LLMs) in use today.
Gradient Heterogeneity Complements Hessian Heterogeneity in Transformer Optimization
The paper investigates why adaptive optimizers like Adam outperform SGD when fine‑tuning Transformers. It introduces gradient heterogeneity—the variation in gradient norms across parameter blocks—and shows, both theoretically and experimentally, that this heterogeneity, together with Hessian heterogeneity, hampers SGD convergence while sign‑based methods such as SignSGD are less affected. The study links the source of gradient heterogeneity to layer‑normalization placement, finding that Post‑LN architectures exhibit the strongest effect, and uses SignSGD as a tractable proxy to analyze Adam‑like behavior and learning‑rate scaling.
Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.