arXiv Machine Learning By Ryan Lee, Jacob Biloki, Edward J. Hu, Jonathan May

Sparse Layers are Critical to Scaling Looped Language Models

Read the original on arXiv Machine Learning →

arXiv:2605. 09165v2 Announce Type: replace Abstract: Looped language models repeat a set of transformer layers through depth, reducing memory costs and providing natural early-exit points at loop boundaries.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv AI
2d ago

Looping Beyond Twice: A Scalable Recipe for Looped Mixture-of-Experts

The paper introduces LOOM, a method for scaling looped mixture‑of‑experts (MoE) Transformers beyond the typical two‑loop limit. LOOM addresses two key obstacles: it stabilizes deep recurrence by bounding residual variance and re‑injecting the input embedding, and it prevents expert selection collapse by using per‑loop routers and a looping residual to maintain computational diversity. Experiments on 100 M–1.7 B parameter models show stable scaling to 9–12 loops, with significant perplexity reductions and zero‑shot accuracy gains under near‑iso‑FLOP conditions.

By Di He, Pengxiang Li, Da Chang, Qingyan Meng, Lu Yin, Shiwei Liu