arXiv Machine Learning

Sparse Layers are Critical to Scaling Looped Language Models

arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.

arXiv AI
2d ago

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.

By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
arXiv Machine Learning
Jul 3

Hyperloop Transformers

arXiv:2604. 21254v3 Announce Type: replace Abstract: LLM architecture research generally aims to maximize model quality subject to fixed compute/latency budgets.

By Abbas Zeitoun, Lucas Torroba-Hennigen, Yoon Kim
arXiv Machine Learning
4d ago

Looped Transformers as Optimizers

arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...

By Yulong Huang, Chen Jiang, Zhanpeng Zhou, Hongtao Zhang, Tianyu Li, Tianyu He, Xiangyu Zhang, Bojun Cheng
arXiv AI
Aug 11

Full-bandwidth transformer

arXiv:2608. 08888v1 Announce Type: new Abstract: Autoregressive transformers compute along two axes: horizontally across generated tokens, and vertically through model depth.

By Xi Wang, Ziyang Cai, Zheng Zhan, Harry Dong, Ying Fan, Gustavo de Rosa, Tim Pearce, John Langford
arXiv Machine Learning
Sep 17

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

The paper demonstrates that architectural changes—specifically looped transformers and boundary operators—can alter scaling exponents in pre‑training, yielding exponential performance gains for a given computational budget. Looping, or recursive depth, enables model growth that matches larger models (e.g., a 7.4B looped architecture matching GPT‑3 13B) with significantly less compute, while boundary operators provide additional, though smaller, efficiency improvements. In data‑constrained, multi‑epoch scenarios, increasing loops with scale serves as a useful regularizer, suggesting that deeper computational depth drives compute‑efficiency gains that grow with model size.

By Zixi Chen, Akshay Vegesna, Samip Dahal, Andrew Gordon Wilson