arXiv Machine Learning

Dense Supervision Is Not Enough: The Readout Blind Spot in Looped Language Models

arXiv:2606. 24898v1 Announce Type: new Abstract: Looped language models turn hidden states into runtime state: each state is decoded for prediction and fed back into future computation.

arXiv AI
Jul 16

DeepLoop: Depth Scaling for Looped Transformers

arXiv:2607. 13491v1 Announce Type: cross Abstract: Looped Transformers scale sequential computation by applying a compact stack of physical blocks for multiple rounds, increasing unrolled depth without increasing stored parameters.

By Shuzhen Li, Yifan Zhang, Jiacheng Guo, Quanquan Gu, Mengdi Wang
arXiv Machine Learning
1d ago

Decoding Looped Transformers Better for (Almost) Free

The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.

By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
arXiv Computation and Language
Sep 16

Persistent Recurrent Memory Between Transformer Layers - Improves Language Model Generalization

The paper proposes a lightweight recurrent memory module inserted between the lower and upper halves of a 6‑layer decoder‑only transformer. This module, which uses cross‑attention to observe hidden states, a GRU to update a persistent state, and gated addition to modulate subsequent layers, adds only 3.7% more parameters. It reduces evaluation loss by 28.5% and narrows the generalization gap, with ablations showing the benefit comes solely from the memory topology rather than auxiliary losses.

By Eduardo Novaes Hering
arXiv Computer Vision
Sep 11

LoopVAE: Recurrent Depth Across Scales for Visual Tokenization

LoopVAE introduces a recurrent depth architecture that reuses a scale‑ and loop‑conditioned core across different spatial scales while keeping resolution‑changing transitions separate. The four‑block core applies 28 block operations per encoder or decoder, enabling a 29M‑parameter convolutional model to achieve 0.28 rFID and 32.54 dB PSNR on ImageNet‑256 with roughly 65% fewer parameters than comparable VAEs. Experiments with both convolutional and Transformer operators, as well as ablations on parameter sharing, demonstrate competitive image quality metrics and reveal how targeted loop interventions and truncation affect reconstruction quality and computational trade‑offs.

By Zhiying Lu
arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
arXiv AI
Sep 15

Bypass Observation: A Conceptual Design of a Non-Intrusive Layer-Wise Semantic Extraction Architecture

The paper proposes Bypass Observation, a non‑intrusive layer‑wise readout architecture that attaches read‑only observation heads to selected Transformer layers without feeding their outputs back into the backbone. Three variants are explored: a shared language‑model head across layers, layer‑specific heads, and a layer‑ or step‑adaptive head. The authors provide a closed‑form overhead estimate (≈ V/(12d)) and discuss ways to reduce cost, while distinguishing bypass chain‑of‑thought from conventional chain‑of‑thought and outlining potential applications to looped and recurrent‑depth Transformers.

By Haibin Tong, Jiang Yu
arXiv AI
Sep 18

Beyond Depth Truncation: Controlled Evaluation of Depth Utilization in Recursive Language Models

The paper critiques the common practice of evaluating depth usage in depth‑recurrent language models by truncating depth during inference and measuring performance decline. It argues that this method conflates three distinct effects—fewer block applications, reduced computation, and an out‑of‑distribution readout—yet is usually interpreted as measuring only the second. To address this, the authors introduce the Depth Control Protocol (DCP), a suite of positive and negative controls that isolate each factor, along with a training intervention to confirm causality, specifically tailored for depth‑wise weight‑sharing architectures.

By Ha Van Dau, Thanh Tung Khuat, Nguyen Thanh Dung