arXiv AI

Stabilizing Extrapolation in Looped Transformers via Learned Stochastic Stopping

arXiv:2606. 29983v1 Announce Type: cross Abstract: Looped Transformers, which repeatedly apply a shared transformer block, are an architecturally natural fit for variable-length algorithmic tasks.

arXiv Machine Learning
Sep 25

FlashLoop: Fast and Memory-Efficient Looped Transformers via Lazy Updates

FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.

By Wanqi Yang, Shiwei Liu
arXiv Machine Learning
4d ago

Decoding Looped Transformers Better for (Almost) Free

The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.

By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
arXiv AI
5d ago

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.

By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu
arXiv AI
2d ago

The Surprising Effectiveness of Shared Memory in Looped Transformers

The paper introduces a method for looped transformers that share a key‑value cache across recursions, reducing memory usage without sacrificing quality. Experiments show that the Looped Prediction Transformer (LPT) and its hybrid variant achieve lower perplexity on FineWeb‑Edu while using 76‑79% less context memory compared to standard transformers. Analysis reveals that shared memory develops distinct representations and serves as a gradient highway, enabling later recursions to focus on shared information.

By Giovanni Monea, Keshav Ramji, Yousef El-Kurdi, Luis A. Lastras, Yoav Artzi, Nathan Godey, Ram\'on Fernandez Astudillo
arXiv Machine Learning
Sep 30

Looped Transformers as Optimizers

arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...

By Yulong Huang, Chen Jiang, Zhanpeng Zhou, Hongtao Zhang, Tianyu Li, Tianyu He, Xiangyu Zhang, Bojun Cheng
arXiv Machine Learning
Sep 30

Scheduling Recursive Reasoning in Looped Transformers

arXiv:2609.36653v1 Announce Type: new Abstract: Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with sha...

By Boyuan Wang, Chengyao Yu, Jiaxi Ren, Hongxin Wei, Bingyi Jing, Yuxin Tao
arXiv Machine Learning
Sep 24

Rolling Conformal Prediction in Sequential Model Training

Rolling Conformal Prediction (rolling‑CP) is a distribution‑free predictive inference method designed for sequential model training. It calibrates each incoming observation against the current predictor and incorporates it into future training, eliminating the need for data splitting. For exchangeable data, rolling‑CP guarantees marginal coverage with a universal factor‑two bound, and for i.i.d. streams it provides high‑probability training‑conditional validity over time, improving to the target level under stability conditions.

By Chen Cheng, Ruiting Liang, Rina Foygel Barber