arXiv AI By Zhiheng Hu, Yixun Wei, Jian Zhou, Yizhuang Zhou, Ji Li, Xing Chen, Yang Li, Bojun Wang, Yibo Zhu, Xiangyu Zhang, Daxin Jiang

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

Read the original on arXiv AI →

KITE (KV-Invariant Transformer Expansion) is a scaling paradigm that trains a language model from a smaller size to a larger one, saving training costs by upcycling. It places new parameters in regions that do not affect attention KV, so inference only requires prefilling KV from the smaller part, reducing inference costs. The Step Scale Transformer (SST), a two-tower decoder, demonstrates this by achieving lower training loss than comparable MoE Transformers while cutting estimated inference cost by 6.7% and 31.6%.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.