arXiv AI

State Propagation Also Satisfies: A Complex-Valued State-Space Model for Deterministic State Tracking

arXiv:2608. 03425v1 Announce Type: new Abstract: Transformer-based architectures have dominated sequence modeling, largely due to the expressive power of attention mechanisms.

arXiv Machine Learning
Jun 26

Learning State-Tracking from Code Using Linear RNNs

arXiv:2602. 14814v3 Announce Type: replace Abstract: Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers and RNNs (linear and non-linear).

By Julien Siems, Riccardo Grazzi, Korbinian P\"oppel, Kirill Kalinin, Hitesh Ballani, Babak Rahmani
arXiv Machine Learning
Aug 31

InfoMamba: An Attention-Free Hybrid Mamba-Transformer Model

InfoMamba is an attention‑free hybrid model that combines a minimal‑bandwidth global interface with a selective recurrent stream. The architecture replaces token‑level self‑attention with a concept bottleneck linear filtering layer and integrates it via an information‑maximizing fusion (IMF) that injects global context into the state‑space dynamics. Experiments across classification, dense prediction, and non‑vision tasks show that InfoMamba outperforms strong Transformer and SSM baselines while maintaining near‑linear scaling and competitive accuracy‑efficiency trade‑offs.

By Youjin Wang, Jiaqiao Zhao, Rong Fu, Run Zhou, Ruizhe Zhang, Jiani Liang, Suisuai Cao, Feng Zhou
arXiv Machine Learning
Jun 11

Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking

arXiv:2602. 10743v2 Announce Type: replace Abstract: State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking.

By Vaisakh Shaj, Cameron Barker, Aidan Scannell, Andras Szecsenyi, Elliot J. Crowley, Amos Storkey
arXiv Machine Learning
Aug 27

Gated Recurrent Transformers: Expressive Depth through Recurrent Modulation

The paper introduces Gated Recurrent Transformers, a depth‑sharing architecture that brackets a single shared core with fixed prelude and coda blocks and uses a lightweight projection and element‑wise update gate to modulate recurrent updates. This design allows functional specialization across recurrences while reducing memory footprint. Experiments show that, under equal FLOPs or parameter budgets, the recurrent model matches or surpasses deeper GPT‑2 Small baselines, achieving similar or better accuracy with fewer parameters and lower peak decoding memory.

By Amr Hegazy, Amr Alanwar, Mostafa Elhoushi
arXiv Machine Learning
Sep 24

Attention Routing Stabilizes Early: Working-Set Inference for Recurrent Language Models

The paper investigates how attention dynamics evolve across recurrent depth in language models, finding that attention support stabilizes early while hidden states and outputs take longer. It proposes WISE, a training‑free method that uses full attention in early steps and then reuses the discovered sparse working set for later steps, preserving performance on multi‑hop QA tasks. Experiments show that WISE maintains quality up to 2K context, offers measurable speedups, and highlights the importance of recurrent discovery of attention support.

By Ke Wan, Chen Chen
arXiv Machine Learning
Aug 17

The Expressive Limits of Diagonal SSMs for State-Tracking

arXiv:2603. 01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable.

By Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath Chandar
arXiv AI
Jun 29

The Context-Ready Transformer

arXiv:2606. 27538v1 Announce Type: cross Abstract: We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters the block.

By Mahesh Godavarti