arXiv Machine Learning

Looped Transformers as Optimizers

arXiv Machine Learning
6d ago

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.

By Wanqi Yang, Shiwei Liu
arXiv AI
1d ago

LoopICL: Looping a single transformer block to solve tabular tasks

LoopICL is a transformer architecture that loops a single block to address tabular tasks. It separates parameter count from computational depth by using a cell stream for per‑cell features and a row stream for in‑context examples, refined via within‑column and cross‑column attention. During pre‑training, varying loop counts and a learned exit‑gate allow the model to adjust inference depth at test time, achieving competitive performance with TabICLv2 while using about 90% fewer parameters.

By Amir Rezaei Balef, Katharina Eggensperger
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
arXiv Machine Learning
Sep 17

How Model Growth, Recursion, and Boundary Operators Influence Scaling Exponents

The paper demonstrates that architectural changes—specifically looped transformers and boundary operators—can alter scaling exponents in pre‑training, yielding exponential performance gains for a given computational budget. Looping, or recursive depth, enables model growth that matches larger models (e.g., a 7.4B looped architecture matching GPT‑3 13B) with significantly less compute, while boundary operators provide additional, though smaller, efficiency improvements. In data‑constrained, multi‑epoch scenarios, increasing loops with scale serves as a useful regularizer, suggesting that deeper computational depth drives compute‑efficiency gains that grow with model size.

By Zixi Chen, Akshay Vegesna, Samip Dahal, Andrew Gordon Wilson
arXiv AI
Jun 17

LoopCoder-v2: Only Loop Once for Efficient Test-Time Computation Scaling

arXiv:2606. 18023v1 Announce Type: cross Abstract: Looped Transformers scale latent computation by repeatedly applying shared blocks, but sequential looping increases latency and KV-cache memory with the loop count.

By Jian Yang, Shawn Guo, Wei Zhang, Tianyu Zheng, Yaxin Du, Haau-Sing Li, Jiajun Wu, Yue Song, Yan Xing, Qingsong Cai, Zelong Huang, Chuan Hao, Ran Tao, Xianglong Liu, Wayne Xin Zhao, Mingjie Tang, Weifeng Lv, Ming Zhou, Bryan Dai