arXiv Machine Learning By Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath Chandar

The Expressive Limits of Diagonal SSMs for State-Tracking

Read the original on arXiv Machine Learning →

arXiv:2603. 01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Statistics ML
2d ago

Extending SSMs with the Exponentially Weighted Signature

arXiv:2603.19198v3 Announce Type: replace Abstract: We introduce the exponentially weighted signature (EWS), a continuous-time model that computes iterated integrals of a path, where each increment i...

By Alexandre Bloch, Benjamin Walker, Jo\"el Mouterde, Sam Morley, Samuel N. Cohen, Terry Lyons
arXiv Machine Learning
Jun 11

Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking

arXiv:2602. 10743v2 Announce Type: replace Abstract: State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking.

By Vaisakh Shaj, Cameron Barker, Aidan Scannell, Andras Szecsenyi, Elliot J. Crowley, Amos Storkey
arXiv AI
Sep 25

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

The paper investigates how neural networks, particularly Transformers and recurrent models, learn to track group elements by predicting their running product. It finds that Transformers tend to recover quotient classes and that their accuracy can be predicted by the reciprocal of class size, while recurrent networks can capture both normal and non‑normal right‑coset partitions. The study links partial accuracy, learning stages, and internal state representations to the subgroup cosets the models learn to track.

By Zhiyu Zhang, Yupeng Li