Scaling Laws for Looped Mixture of Experts
arXiv:2609.40316v1 Announce Type: cross Abstract: Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed p...
arXiv:2609.40316v1 Announce Type: cross Abstract: Looped transformers and Mixture-of-Experts (MoE) offer complementary routes to efficient scaling: recurrence increases computational depth at fixed p...
arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.
arXiv:2609.01343v1 Announce Type: new Abstract: Looped Transformers increase effective depth by iterating a shared block of layers, but most evaluations compare at fixed model size, conflating archit...
The paper demonstrates that architectural changes—specifically looped transformers and boundary operators—can alter scaling exponents in pre‑training, yielding exponential performance gains for a given computational budget. Looping, or recursive depth, enables model growth that matches larger models (e.g., a 7.4B looped architecture matching GPT‑3 13B) with significantly less compute, while boundary operators provide additional, though smaller, efficiency improvements. In data‑constrained, multi‑epoch scenarios, increasing loops with scale serves as a useful regularizer, suggesting that deeper computational depth drives compute‑efficiency gains that grow with model size.
arXiv:2606. 04438v1 Announce Type: cross Abstract: Mixture-of-Experts (MoE) and looped architectures scale models along two orthogonal axes, namely parameter capacity and effective depth.
The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.
arXiv:2609.37379v1 Announce Type: new Abstract: Looped Transformers provide a parameter-efficient approach to depth scaling by repeatedly applying shared Transformer blocks. Recent reasoning models h...
arXiv:2607. 22577v1 Announce Type: new Abstract: Scaling large language models (LLMs) has driven their success, yet dense Transformers couple capacity and computation: every parameter is activated for every token, making training and inference costs grow linearly with model size-a critical bottleneck as models approach trillion-parameter regimes.
FlashLoop is a training‑free inference framework for Looped Transformers that reduces cross‑loop redundancy by employing token‑sparse updates, sparse attention, and KV‑residual quantization. It exploits observations that, as loops progress, state changes concentrate on a small token subset, attention differences are dominated by a sparse key subset, and KV residuals become amenable to low‑bit quantization. The method achieves lossless accuracy with up to 1.64× speedup and 6× KV‑cache memory reduction across several Looped Transformer models.
arXiv:2609.36636v1 Announce Type: new Abstract: Looped language models (LoopLMs) increase computational depth through parameter sharing, offering a path to scale inference computation without adding...
The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.
Scaling Large Language Models (LLMs) has been driven mainly by enlarging the Transformer backbone, but for an already-strong model this requires another round of costly pretraining. We study whether an existing backbone can keep improving by allocating more computation to each token while leaving the Transformer backbone fixed.