How we sped up transformer inference 100x for 🤗 API customers
Related stories
Our Transformers Code Agent beats the GAIA benchmark 🏅
KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling
KITE (KV-Invariant Transformer Expansion) is a scaling paradigm that trains a language model from a smaller size to a larger one, saving training costs by upcycling. It places new parameters in regions that do not affect attention KV, so inference only requires prefilling KV from the smaller part, reducing inference costs. The Step Scale Transformer (SST), a two-tower decoder, demonstrates this by achieving lower training loss than comparable MoE Transformers while cutting estimated inference cost by 6.7% and 31.6%.
Performance-Efficiency Tradeoffs in Transformers: An Approximation Theory Perspective
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layers’ role in information extraction and characterizes the trade‑off between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.
License to Call: Introducing Transformers Agents 2.0
Overview of natively supported quantization schemes in 🤗 Transformers
Small-Scale Experiments: Are We There Yet?
arXiv:2608. 11859v1 Announce Type: new Abstract: Scaling laws promised cost-effective experiments; six years later, they have yet to fully deliver.
At the Edge of Understanding: Sparse Autoencoders Trace The Limits of Transformer Generalization
arXiv:2606. 26396v1 Announce Type: new Abstract: Pre-trained transformers have demonstrated remarkable generalization abilities, at times extending beyond the scope of their training data.
Introducing Optimum: The Optimization Toolkit for Transformers at Scale
Tokenization in Transformers v5: Simpler, Clearer, and More Modular
Experimenting with the proposed Cross-Origin Storage API in Transformers.js
Revisiting the Shape Convention of Transformer Language Models
arXiv:2602.06471v2 Announce Type: replace Abstract: The architectural shape of dense Transformers has remained remarkably stable: narrow-wide-narrow feed-forward networks (FFNs) consume most non-embe...