arXiv:2609.38356v1 Announce Type: new
Abstract: Dynamical Systems Reconstruction (DSR) aims to infer models from observed time series that reproduce a system's qualitative long-term behavior. Continu...
By Sima Hashemi, Daniel Durstewitz, Georgia Koppe
arXiv:2607. 27656v1 Announce Type: new Abstract: Looped Transformers create a useful train- and test-time compute axis by reusing the same Transformer block over recurrent depth, increasing effective depth at a fixed parameter count.
By Bum Jun Kim, Kohei Hayashi, Shunsuke Kamiya, Masanori Koyama, Yusuke Iwasawa, Yutaka Matsuo
arXiv:2602. 18131v2 Announce Type: replace Abstract: Temporal Predictive Coding provides a layer-local, parallelisable mechanism for learning in recurrent systems, making it an attractive candidate for online local learning on neuromorphic and edge hardware.
By Tom Potter, Oliver Rhodes
Dynamic Compression in Recurrent Networks proposes a method that lets recurrent models revisit and revise their fixed-size state through additional updates, rather than compressing all information in a single causal pass. This approach allows the model to retain lower-fidelity history and refine only the relevant parts when needed, reducing the required state size for accurate task reuse. Experiments show that dynamic compression lowers the recurrent state needed and scales better as the number of stored functions increases.
By Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
By Alessandro Ronca
The paper introduces ELiSe, a model that leverages cortical network scaffolds and dendritic compartments to learn complex non‑Markovian spatio‑temporal patterns using only local, always‑on, phase‑free synaptic plasticity. It demonstrates the model’s ability to acquire and replay intricate sequences, exemplified by a birdsong learning mock‑up, and shows robustness to external disturbances and flexibility in parameter settings.
By Laura Kriener, Kristin V\"olk, Ben von H\"unerbein, Federico Benitez, Walter Senn, Mihai A. Petrovici
arXiv:2610. 01369v1 Announce Type: new Abstract: Understanding a nonlinear dynamical system from time series requires not only reproducing its trajectories, but also identifying a simple representation that preserves its essential dynamical structure.
By Hiroto Tamura, Gouhei Tanaka
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
By Shota Takashiro, Masanori Koyama, Takeru Miyato, Yusuke Iwasawa, Yutaka Matsuo, Kohei Hayashi
arXiv:2608. 11690v1 Announce Type: new Abstract: Continual learning must absorb new tasks without erasing old ones, and replay---mixing a small buffer of past examples into current training---is among the most effective remedies for catastrophic forgetting.
By Tieliang Gong, Zhongbo Zhang, Wen Wen, Yong-Jin Liu
The paper investigates how pre‑training on random input‑output mappings (memorization tasks) can transfer to downstream tasks. It discovers two unexpected patterns: equivalent transfer, where each pre‑training epoch saves roughly one fine‑tuning epoch, and non‑equivalent transfer, where pre‑training on a mismatched task can be more efficient than training directly on the downstream task. Ablation studies reveal that transfer consists of a trivial magnitude‑driven effect in the last layer and a non‑trivial structure‑driven effect linked to covariance in other layers.
By Yimiao Yu, Florentin Guth
The paper introduces Recursive Quadrature Filters (RQFs), complex‑valued temporal filters that act as band‑pass filters within diagonal state‑space models. By making each layer’s bottom‑up input prospective through a parameter‑free two‑tap update, the authors mitigate depth‑dependent gradient attenuation in deep continuous‑time recurrent networks. Experiments on RQFs, S5, and ORGaNICs show that prospective variants match or surpass non‑prospective controls, achieving high accuracy on raw‑audio Speech Commands and the Path‑X task with few parameters.
By Shivang Rawat, Mirko Morello, Flaviano Morone, David J. Heeger
arXiv:2608. 15062v1 Announce Type: cross Abstract: Scaling transformer language models creates an inherent tension between expressivity and memory efficiency.
By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi