arXiv Machine Learning

A Held-Out Transition-Pair Falsifier for Long-Horizon Non-Abelian State Tracking

arXiv:2606. 07254v1 Announce Type: new Abstract: State tracking exposes a sharp limitation of sequence models: the relevant signal is often not a summary of observed tokens, but an ordered latent state that evolves through non-commutative transformations.

arXiv Machine Learning
Aug 17

The Expressive Limits of Diagonal SSMs for State-Tracking

arXiv:2603. 01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable.

By Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath Chandar
arXiv AI
Sep 25

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

The paper investigates how neural networks, particularly Transformers and recurrent models, learn to track group elements by predicting their running product. It finds that Transformers tend to recover quotient classes and that their accuracy can be predicted by the reciprocal of class size, while recurrent networks can capture both normal and non‑normal right‑coset partitions. The study links partial accuracy, learning stages, and internal state representations to the subgroup cosets the models learn to track.

By Zhiyu Zhang, Yupeng Li
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa
arXiv Machine Learning
1d ago

Decoding Looped Transformers Better for (Almost) Free

The paper introduces LoopCD, a training‑free contrastive decoding framework that improves token selection in Loop‑Transformer models by comparing the final prediction with earlier recurrent passes. LoopCD operates either in logit space (LoopCD‑Logits) with a single extra output pass or in hidden‑state space (LoopCD‑Hidden) with no output overhead. Across multiple looped Transformer families, LoopCD yields significant performance gains—raising pass@1 scores on tasks such as AIME 2024 and HumanEval—while enabling a reduction in the number of recurrent loops and a corresponding decrease in inference FLOPs.

By Weihao Liu, Huangjie Zheng, Tianrong Chen, Rohit Dilip, Richard He Bai, Yizhu Jiao, Yuyang Wang, Ruixiang Zhang
arXiv Machine Learning
Jun 26

Learning State-Tracking from Code Using Linear RNNs

arXiv:2602. 14814v3 Announce Type: replace Abstract: Over the last years, state-tracking tasks, particularly permutation composition, have become a testbed to understand the limits of sequence models architectures like Transformers and RNNs (linear and non-linear).

By Julien Siems, Riccardo Grazzi, Korbinian P\"oppel, Kirill Kalinin, Hitesh Ballani, Babak Rahmani
arXiv Machine Learning
Aug 11

Advancing Intelligent Sequence Modeling: Evolution, Trade-offs, and Applications of State-Space Architectures from S4 to Mamba

arXiv:2503. 18970v4 Announce Type: replace Abstract: Structured State Space Models (SSMs) have become a prominent class of sequence models, developed against two long-standing difficulties: the sequential computation and gradient propagation limits of Recurrent Neural Networks (RNNs), and the quadratic time and memory cost of self-attention in Transformers.

By Shriyank Somvanshi, Md Monzurul Islam, Mahmuda Sultana Mimi, Sazzad Bin Bashar Polock, Gaurab Chhetri, Anandi Dutta, Amir Rafe, Subasish Das