Learning to Remember: Distilling Memory Retention for Compact Recurrent Neural Networks
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 06796v1 Announce Type: cross Abstract: Deep learning has achieved remarkable success in various domains including time series analysis, computer vision and natural language processing.
arXiv:2606.21562v2 Announce Type: replace Abstract: Transformers are AI's workhorse but their computational cost becomes prohibitive when processing long sequences. We target long-horizon streaming v...
The paper introduces the Continuous Memory Machine (CMM), a recurrent neural network that uses separate matrix-valued short‑term and long‑term memory states. Short‑term memory tracks recent neural activity with neuron‑level models for rapid computation, while long‑term memory stores information for later use; both are jointly updated by a Transformer that allows bidirectional read‑write operations. Experiments on algorithmic, in‑context learning, and recurrent reasoning tasks show that CMM outperforms many baselines and generalizes better than previous memory‑augmented networks, while maintaining interpretable attention patterns from the Continuous Thought Machine.
arXiv:2609.06006v1 Announce Type: cross Abstract: Deep learning for time series has progressed through successive architectural paradigms, from recurrent networks and transformers to structured state...
arXiv:2511. 20577v5 Announce Type: replace Abstract: Real-world time series often exhibit strong non-stationarity, complex nonlinear dynamics, and behavior expressed across multiple temporal scales, from rapid local fluctuations to slow-evolving long-range trends.
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.