arXiv Machine Learning

Shared Weights, Selected Computations: How Looped Transformers Route What Each Loop Does

arXiv AI
Sep 3

Looped Transformers under the Jacobian Lens: Does the Global Workspace Survive Recurrence?

The paper investigates whether a global workspace—a set of verbalisable, causally potent representations—emerges in transformer models that use recurrence instead of a stack of distinct layers. Using a Jacobian lens extended with a virtual‑unrolling adapter, the authors analyze two recurrent transformer architectures, Ouro‑2.6B and Huginn‑0125, and compare them to a standard Qwen3.6‑27B baseline. They find that a workspace does form in the iterated parts of both models, but recurrence alters how it can be accessed: Ouro reconstructs workspace content in every loop and requires writes and ablations across all loops, whereas Huginn forwards content across all recurrences but limits reads, writes, and ablations to a sliding window of about two recurrences.

By Wenlong Wang, Fergal Reid
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
arXiv Machine Learning
Jul 2

Is One Layer Enough? Training A Single Transformer Layer Can Match Full-Parameter RL Training

arXiv:2607. 01232v1 Announce Type: new Abstract: Reinforcement learning (RL) has become a central component of post-training large language models (LLMs), yet little is understood about how RL adaptation is distributed across transformer layers.

By Zijian Zhang, Rizhen Hu, Athanasios Glentis, Dawei Li, Chung-Yiu Yau, Hongzhou Lin, Mingyi Hong