arXiv AI

Continuous Memory Machines

The paper introduces the Continuous Memory Machine (CMM), a recurrent neural network that uses separate matrix-valued short‑term and long‑term memory states. Short‑term memory tracks recent neural activity with neuron‑level models for rapid computation, while long‑term memory stores information for later use; both are jointly updated by a Transformer that allows bidirectional read‑write operations. Experiments on algorithmic, in‑context learning, and recurrent reasoning tasks show that CMM outperforms many baselines and generalizes better than previous memory‑augmented networks, while maintaining interpretable attention patterns from the Continuous Thought Machine.

arXiv Machine Learning
Aug 19

Dynamic Compression in Recurrent Networks

Dynamic Compression in Recurrent Networks proposes a method that lets recurrent models revisit and revise their fixed-size state through additional updates, rather than compressing all information in a single causal pass. This approach allows the model to retain lower-fidelity history and refine only the relevant parts when needed, reducing the required state size for accurate task reuse. Experiments show that dynamic compression lowers the recurrent state needed and scales better as the number of stored functions increases.

By Jyothish Pari, Ryan Bahlous-Boldi, Pulkit Agrawal
arXiv Machine Learning
Sep 29

Progressive Memory Transformer: Memory-Aware Attention for Time-Series

The paper introduces the Progressive Memory Transformer (PMT), a transformer variant that adds writable, window‑aligned memory to expose mid‑range representations alongside token and sequence‑level outputs. PMT is trained with a hierarchical learning framework that applies separate objectives at local, mid‑range, and global scales, encouraging the model to capture fine‑grained variation, window‑level motifs, and overall sequence agreement. Experiments on seven UCR/UEA/UCI classification datasets, a cue‑retention probe, and forecasting tasks show that PMT achieves strong low‑label classification performance, competitive multi‑horizon forecasting, and evidence that its memory states encode mid‑range motifs.

By Tord Sture Stangeland, Andreas K\"ohler, Steffen M{\ae}land, Ad\'in Ram\'ires Rivera