arXiv Machine Learning By Zixi Chen, Akshay Vegesna, Samip Dahal, Andrew Gordon Wilson

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

Read the original on arXiv Machine Learning →

The paper demonstrates that architectural changes—specifically looped transformers and boundary operators—can alter scaling exponents in pre‑training, yielding exponential performance gains for a given computational budget. Looping, or recursive depth, enables model growth that matches larger models (e.g., a 7.4B looped architecture matching GPT‑3 13B) with significantly less compute, while boundary operators provide additional, though smaller, efficiency improvements. In data‑constrained, multi‑epoch scenarios, increasing loops with scale serves as a useful regularizer, suggesting that deeper computational depth drives compute‑efficiency gains that grow with model size.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
Jul 16

DeepLoop: Depth Scaling for Looped Transformers

arXiv:2607. 13491v1 Announce Type: cross Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters.

By Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
arXiv AI
Sep 24

KITE: KV-Invariant Transformer Expansion for Efficient Agentic LLM Scaling

KITE (KV-Invariant Transformer Expansion) is a scaling paradigm that trains a language model from a smaller size to a larger one, saving training costs by upcycling. It places new parameters in regions that do not affect attention KV, so inference only requires prefilling KV from the smaller part, reducing inference costs. The Step Scale Transformer (SST), a two-tower decoder, demonstrates this by achieving lower training loss than comparable MoE Transformers while cutting estimated inference cost by 6.7% and 31.6%.

By Zhiheng Hu, Yixun Wei, Jian Zhou, Yizhuang Zhou, Ji Li, Xing Chen, Yang Li, Bojun Wang, Yibo Zhu, Xiangyu Zhang, Daxin Jiang
arXiv Machine Learning
1d ago

Looped Transformers as Optimizers

arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...

By Yulong Huang, Chen Jiang, Zhanpeng Zhou, Hongtao Zhang, Tianyu Li, Tianyu He, Xiangyu Zhang, Bojun Cheng