arXiv:2606. 00732v1 Announce Type: new Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings.
By Jayanta Dey, Shikhar Srivastava, Itamar Lerner, Christopher Kanan, Dhireesha Kudithipudi
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
By Akarsh Kumar, Phillip Isola
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
By Johannes von Oswald, Nino Scherrer, Seijin Kobayashi, Luca Versari, Songlin Yang, Sarthak Mittal, Maximilian Schlegel, Kaitlin Maile, Yanick Schimpf, Oliver Sieberling, Alexander Meulemans, Rif A. Saurous, Guillaume Lajoie, Charlotte Frenkel, Razvan Pascanu, Blaise Ag\"uera y Arcas, Jo\~ao Sacramento
arXiv:2410. 11687v3 Announce Type: replace-cross Abstract: Linear recurrent networks (LRNNs) offer linear-time sequence modeling, but standard recurrent updates do not directly expose the supervised products needed for in-context gradient descent.
By Yudou Tian, Neeraj Mohan Sushma, Harshvardhan Mestha, Nicolo Colombo, David Kappel, Anand Subramoney
The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
By Ke Wan, Chen Chen
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.
By Jayden Teoh, Manan Tomar, Kwangjun Ahn, Edward S. Hu, Tim Pearce, Pratyusha Sharma, Akshay Krishnamurthy, Riashat Islam, Alex Lamb, John Langford
arXiv:2606. 24969v1 Announce Type: new Abstract: While the quadratic sequence-length bottleneck of transformers has fueled a resurgence in recurrent models, effectively capturing complex dynamics requires architectures that balance efficient training with highly expressive latent states.
By Klaus Schertler, Xiomara Runge, Andrea Ceni, David Kappel, Claudio Gallicchio
Fast Weight Attention for Continual Learning introduces recurrent fast‑weight memories and selective state‑space models that compress expanding context into a fixed‑size recurrent state, enabling an online learning rule for state transitions. The paper derives normalized first‑order updates for squared‑error regression and negative inner‑product objectives, presenting several variants (Falcon‑1, Falcon‑2, Falcon‑3 and their inner‑product counterparts) with recurrent, masked‑parallel, and chunk‑parallel implementations. These methods demonstrate competitive performance in language modeling and improved length extrapolation on variable‑digit addition tasks.
By Yifan Zhang, Steve Ta, Jasper Zhang, Jichen Feng, Shuzhen Li, Yongxin Zhang, Yifeng Liu, Huizhuo Yuan, Mengdi Wang, Quanquan Gu, Andrew Chi-Chih Yao
arXiv:2609.16540v1 Announce Type: cross
Abstract: State Space Models (SSMs) have emerged as a compelling alternative to Transformers, enabling sequence modeling with constant memory and linear comput...
By William L. Tong, Aryo Lotfi, Emmanuel Abbe, Kostas Vaggelakos, Vishnu Banna, Etai Littwin, Josh Susskind, Cengiz Pehlevan, Eran Malach
arXiv:2606.21562v2 Announce Type: replace
Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
By Philippe Weinzaepfel, Christian Wolf, Mert B\"ulent Sariyildiz, Guillaume Bono, Gianluca Monaci
The paper introduces the Progressive Memory Transformer (PMT), a transformer variant that adds writable, window‑aligned memory to expose mid‑range representations alongside token and sequence‑level outputs. PMT is trained with a hierarchical learning framework that applies separate objectives at local, mid‑range, and global scales, encouraging the model to capture fine‑grained variation, window‑level motifs, and overall sequence agreement. Experiments on seven UCR/UEA/UCI classification datasets, a cue‑retention probe, and forecasting tasks show that PMT achieves strong low‑label classification performance, competitive multi‑horizon forecasting, and evidence that its memory states encode mid‑range motifs.
By Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera
The paper investigates how temporal recurrence affects the required depth of neural networks in streaming tasks. By treating depth, expert width, and parallel experts as a compute‑allocation problem, the authors compare recurrent and non‑recurrent models across various compute budgets. Experiments on Sokoban and FineWeb language modeling show that recurrence shifts the optimal compute allocation toward fewer layers while maintaining or improving performance.
By Ivan Anokhin, Johan Obando-Ceron, Irina Rish, Sebastian Risi