arXiv:2609.09157v1 Announce Type: new
Abstract: Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their t...
By Hanwen Jiang
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
By Akarsh Kumar, Phillip Isola
Dynamic Compression in Recurrent Networks proposes a method that lets recurrent models revisit and revise their fixed-size state through additional updates, rather than compressing all information in a single causal pass. This approach allows the model to retain lower-fidelity history and refine only the relevant parts when needed, reducing the required state size for accurate task reuse. Experiments show that dynamic compression lowers the recurrent state needed and scales better as the number of stored functions increases.
By Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal
Dynamic Compression in Recurrent Networks proposes a method for recurrent models to selectively revisit and update past tokens, rather than compressing all history in a single causal pass. By allowing the model to refine its fixed-size state only when needed, it can maintain lower-fidelity information in the raw sequence and revisit it later. Experiments show that this selective re-scanning reduces the recurrent state needed for accurate task reuse and scales better as the number of stored functions increases.
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
By Shota Takashiro, Masanori Koyama, Takeru Miyato, Yusuke Iwasawa, Yutaka Matsuo, Kohei Hayashi
arXiv:2608. 07420v1 Announce Type: new Abstract: World models are expected to support imagination over extended temporal horizons, yet most are still trained through local few-step prediction objectives and deployed by recursively rolling out their own predictions.
By Xinyi Li, Zaishuo Xia, Chenjie Hao, Yubei Chen
arXiv:2602. 18131v2 Announce Type: replace Abstract: Temporal Predictive Coding provides a layer-local, parallelisable mechanism for learning in recurrent systems, making it an attractive candidate for online local learning on neuromorphic and edge hardware.
By Tom Potter, Oliver Rhodes
arXiv:2609.36337v1 Announce Type: new
Abstract: Tabular foundation models achieve strong performance by conditioning on labelled examples in context, but softmax attention limits their use on large d...
By David Schnurr, Felix Sarnthein, Thomas Hofmann, Imanol Schlag
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.
By Ke Wan, Chen Chen
The paper introduces Recursive Quadrature Filters (RQFs), complex‑valued temporal filters that act as band‑pass filters within diagonal state‑space models. By making each layer’s bottom‑up input prospective through a parameter‑free two‑tap update, the authors mitigate depth‑dependent gradient attenuation in deep continuous‑time recurrent networks. Experiments on RQFs, S5, and ORGaNICs show that prospective variants match or surpass non‑prospective controls, achieving high accuracy on raw‑audio Speech Commands and the Path‑X task with few parameters.
By Shivang Rawat, Mirko Morello, Flaviano Morone, David J. Heeger
arXiv:2606. 00732v1 Announce Type: new Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings.
By Jayanta Dey, Shikhar Srivastava, Itamar Lerner, Christopher Kanan, Dhireesha Kudithipudi