Looped Transformers as Optimizers
arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...
arXiv:2606. 18206v1 Announce Type: new Abstract: Looped architectures provide an inductive bias toward learning step-by-step procedures for tasks that require compositional reasoning.
arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...
arXiv:2604. 11912v2 Announce Type: replace-cross Abstract: While next-token prediction (NTP) has been the standard objective for training language models, it often struggles to capture global structure in reasoning tasks.
arXiv:2607. 15178v1 Announce Type: cross Abstract: Transformer reasoning is limited by autoregressive decoding, which repeat edly compresses rich hidden computation through token space and makes it difficult for intermediate reasoning states to persist across time.
arXiv:2609.39967v1 Announce Type: cross Abstract: Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few...
arXiv:2609.36653v1 Announce Type: new Abstract: Recurrent reasoning models have attracted growing attention for scaling test-time computation, typically by iteratively refining latent states with sha...
arXiv:2607. 19405v1 Announce Type: new Abstract: The CoTFormer architecture formalizes Chain-of-Thought as a form of recurrent latent computation, preserving intermediate states as attendable representations to mimic explicit reasoning traces.
Recursive reasoning models apply a small shared Transformer block many times to refine a latent state. This gives them large effective depth with few parameters and makes them strong on algorithmic ta...
arXiv:2607. 25915v1 Announce Type: new Abstract: Complex structured reasoning tasks often require additional computation, yet current language models obtain it mainly by increasing parameter scale or by serializing intermediate steps as chain-of-thought (CoT) tokens.
The paper introduces looped flows, a new approach that trains looped models using local denoising objectives to overcome the difficulty of training early updates for future ones. By enforcing temporal association through progressively decreasing noise levels and shared noise, the method encourages recurrent states to transfer useful computation over time. Inference is framed as integrating the velocity of a probability flow parameterized by the learned denoiser, allowing the model to solve harder problems by allocating more computation and producing multiple valid predictions from different initial noise samples. Across six reasoning benchmarks, looped flows outperform prior state‑of‑the‑art looped models, achieving 58.8% accuracy on ARC‑AGI‑1 and 12.2% on ARC‑AGI‑2.
T-LoopFormer introduces token-level elastic-depth looped transformers that allow each token to decide its own number of loop iterations based on its hidden state, improving token generation accuracy. It also adds a recursion-wise key‑value cache so tokens at different depths only attend to their corresponding cached states, speeding up autoregressive decoding. Experiments demonstrate strong performance on language modeling and zero‑shot reasoning, achieving the lowest decoding latency among comparable models.
arXiv:2602. 17993v2 Announce Type: replace-cross Abstract: Complex problems, whether in math, logic, or planning, are solved by humans through a sequence of steps where the result of one step informs the next.
arXiv:2606. 07108v1 Announce Type: new Abstract: Recent advances in Large Reasoning Models (LRMs) demonstrate remarkable performance improvements by iteratively reflecting, exploring, and executing complex tasks, yet suffer from inefficiencies due to redundant reasoning, known as "overthinking".