arXiv AI

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

KITE (KV-Invariant Transformer Expansion) is a scaling paradigm that trains a language model from a smaller size to a larger one, saving training costs by upcycling. It places new parameters in regions that do not affect attention KV, so inference only requires prefilling KV from the smaller part, reducing inference costs. The Step Scale Transformer (SST), a two-tower decoder, demonstrates this by achieving lower training loss than comparable MoE Transformers while cutting estimated inference cost by 6.7% and 31.6%.

arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Sep 17

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

The paper demonstrates that architectural changes—specifically looped transformers and boundary operators—can alter scaling exponents in pre‑training, yielding exponential performance gains for a given computational budget. Looping, or recursive depth, enables model growth that matches larger models (e.g., a 7.4B looped architecture matching GPT‑3 13B) with significantly less compute, while boundary operators provide additional, though smaller, efficiency improvements. In data‑constrained, multi‑epoch scenarios, increasing loops with scale serves as a useful regularizer, suggesting that deeper computational depth drives compute‑efficiency gains that grow with model size.

By Zixi Chen, Akshay Vegesna, Samip Dahal, Andrew Gordon Wilson