arXiv Machine Learning

Dynamic Compression in Recurrent Networks

Dynamic Compression in Recurrent Networks proposes a method that lets recurrent models revisit and revise their fixed-size state through additional updates, rather than compressing all information in a single causal pass. This approach allows the model to retain lower-fidelity history and refine only the relevant parts when needed, reducing the required state size for accurate task reuse. Experiments show that dynamic compression lowers the recurrent state needed and scales better as the number of stored functions increases.

Hugging Face Trending Papers
Aug 18

Dynamic Compression in Recurrent Networks

Dynamic Compression in Recurrent Networks proposes a method for recurrent models to selectively revisit and update past tokens, rather than compressing all history in a single causal pass. By allowing the model to refine its fixed-size state only when needed, it can maintain lower-fidelity information in the raw sequence and revisit it later. Experiments show that this selective re-scanning reduces the recurrent state needed for accurate task reuse and scales better as the number of stored functions increases.

arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
arXiv AI
Jun 4

MesaNet: Sequence Modeling by Locally Optimal Test-Time Training

arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.

By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
arXiv AI
Jun 9

MinMax Recurrent Neural Cascades

arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.

By Alessandro Ronca
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen
arXiv Machine Learning
Jun 25

Frequency Domain Reservoir Computing

arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.

By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio