MesaNet: Sequence Modeling by Locally Optimal Test-Time Training
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
arXiv:2606. 06479v1 Announce Type: new Abstract: Training recurrent neural networks (RNNs) requires assigning credit across long sequences of computations.
arXiv:2506. 05233v2 Announce Type: replace-cross Abstract: Sequence modeling is currently dominated by causal transformer architectures that use softmax self-attention.
arXiv:2608. 02870v1 Announce Type: new Abstract: We introduce \ours{}, a recurrent Transformer architecture with fixed-size memory that generalizes sliding-window attention while remaining parallelizable during training.
arXiv:2604. 01577v3 Announce Type: replace-cross Abstract: We study out of distribution generalization in streaming tasks where models are trained on short sequences but must operate over much longer, unknown horizons under bounded memory.
arXiv:2605. 06384v3 Announce Type: replace-cross Abstract: We introduce MinMax Recurrent Neural Cascades (MinMax RNCs), a class of recurrent neural networks built from a novel form of recurrence over the MinMax algebra.
The paper investigates why recurrent models often fail to generalize beyond their training horizon, noting that vanishing or exploding gradients are not the sole cause. It introduces the concept of state credit—the influence of future losses on earlier recurrent states—and proposes Credit Stabilization through Time (CST), a method that rescales this signal during backpropagation to stabilize its norm. Experiments on synthetic and real data show that CST enables models to perform well up to 128 times longer than their training length.
arXiv:2605.26797v2 Announce Type: replace Abstract: We study Latent Recurrent Transformer (LRT), a lightweight augmentation of autoregressive transformers that reuses a high-level source-layer hidden...
arXiv:2410. 11687v3 Announce Type: replace-cross Abstract: Linear recurrent networks (LRNNs) offer linear-time sequence modeling, but standard recurrent updates do not directly expose the supervised products needed for in-context gradient descent.
arXiv:2609.09157v1 Announce Type: new Abstract: Recurrent models provide a natural path to long-context modeling, yet models trained with backpropagation through time (BPTT) often fail beyond their t...
arXiv:2511. 05963v4 Announce Type: replace Abstract: Transformers replace recurrence with a memory that grows with sequence length and self-attention that enables ad-hoc lookups over past tokens.
arXiv:2606. 00732v1 Announce Type: new Abstract: Learning long-range non-stationary temporal patterns remains a core challenge for modern sequence models, particularly in strict streaming settings.
arXiv:2605. 08696v4 Announce Type: replace-cross Abstract: Over the last two decades, language modeling has experienced a shift from the use of predominantly recurrent architectures that process tokens sequentially during training and inference to non-recurrent models that process sequence elements in parallel during training, which results in greater training efficiency and stability at the expense of lower inference throughput.
The paper introduces Decision Titan, a variant of the Decision Transformer that incorporates Test‑Time Training (TTT) layers to store episodic memories in network parameters. It evaluates this architecture on the X‑Maze environment, showing that Decision Titan can learn long‑term dependencies up to 20 times longer than its context window and generalise to sequences 1.7 times longer than the training data. The study also finds that temporal generalisation depends on the choice of time embeddings and that the ability to learn long‑term dependencies hinges on how relevant information is encoded.