FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.
By Wanqi Yang, Shiwei Liu
The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.
By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.
By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
The paper introduces a method for looped transformers that share a key‑value cache across recursions, reducing memory usage without sacrificing quality. Experiments show that the Looped Prediction Transformer (LPT) and its hybrid variant achieve lower perplexity on FineWeb‑Edu while using 76‑79% less context memory compared to standard transformers. Analysis reveals that shared memory develops distinct representations and serves as a gradient highway, enabling later recursions to focus on shared information.
By Giovanni Monea, Keshav Ramji, Yousef El-Kurdi, Luis A. Lastras, Yoav Artzi, Nathan Godey, Ram\'on Fernandez Astudillo
arXiv:2609.37379v1 Announce Type: new
Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...
By Yulong Huang, Chen Jiang, Zhanpeng Zhou, Hongtao Zhang, Tianyu Li, Tianyu He, Xiangyu Zhang, Bojun Cheng
arXiv:2607. 20519v1 Announce Type: new Abstract: Looped Transformers increase test-time computation by repeatedly applying a shared recurrent block.
By Andrei Cristian Popescu, Haitz S\'aez de Oc\'ariz Borde, Pietro Li\`o
Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed parameters, while MoE sparsity expands total capacit...
arXiv:2607. 20594v1 Announce Type: cross Abstract: When does a weight-tied looped transformer -- one block applied T times -- implement an actual algorithm?
By Tong Zhang, Junhao Hu, Yun Peng, Tao Xie
arXiv:2608. 14774v1 Announce Type: new Abstract: Modern sequence models heavily rely on massive memory footprints and large-batch stochastic optimization, barriers that restrict sample efficiency and continual learning.
By Vladimer Khasia
arXiv:2609.36653v1 Announce Type: new
Abstract: Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with sha...
By Boyuan Wang, Chengyao Yu, Jiaxi Ren, Hongxin Wei, Bingyi Jing, Yuxin Tao
arXiv:2605.10335v2 Announce Type: replace-cross
Abstract: Adaptive optimizers such as Adam are standard for training Transformers, but storing gradient first and second moments incurs substantial mem...
By Yao Lu, Dengdong Fan, Shixun Zhang, Yonghong Tian
Rolling Conformal Prediction (rolling‑CP) is a distribution‑free predictive inference method designed for sequential model training. It calibrates each incoming observation against the current predictor and incorporates it into future training, eliminating the need for data splitting. For exchangeable data, rolling‑CP guarantees marginal coverage with a universal factor‑two bound, and for i.i.d. streams it provides high‑probability training‑conditional validity over time, improving to the target level under stability conditions.
By Chen Cheng, Ruiting Liang, Rina Foygel Barber