arXiv Machine Learning
6d ago

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.

By Wanqi Yang, Shiwei Liu
arXiv AI
1d ago

LoopICL: Looping a single transformer block to solve tabular tasks

LoopICL is a transformer architecture that loops a single block to address tabular tasks. It separates parameter count from computational depth by using a cell stream for per‑cell features and a row stream for in‑context examples, refined via within‑column and cross‑column attention. During pre‑training, varying loop counts and a learned exit‑gate allow the model to adjust inference depth at test time, achieving competitive performance with TabICLv2 while using about 90% fewer parameters.

By Amir Rezaei Balef, Katharina Eggensperger