How we sped up transformer inference 100x for π€ API customers
Read the original on Hugging Face Blog βThe Flow has not summarised this story yet β read it at Hugging Face Blog.
The Flow has not summarised this story yet β read it at Hugging Face Blog.
KITE (KV-Invariant Transformer Expansion) is a scaling paradigm that trains a language model from a smaller size to a larger one, saving training costs by upcycling. It places new parameters in regions that do not affect attention KV, so inference only requires prefilling KV from the smaller part, reducing inference costs. The Step Scale Transformer (SST), a two-tower decoder, demonstrates this by achieving lower training loss than comparable MoE Transformers while cutting estimated inference cost by 6.7% and 31.6%.
The paper examines how to allocate attention heads and head dimensions across Transformer layers to balance expressivity and efficiency. It provides a mathematical analysis of early layersβ role in information extraction and characterizes the tradeβoff between head count and dimension under a fixed parameter budget. The authors prove a saturation effect of softmax activations, showing that increasing head dimensions yields diminishing returns, especially for long sequences, and propose strategies for efficient parameter allocation across layers.