arXiv Machine Learning By Jeonghoon Lee

A Held-Out Transition-Pair Falsifier for Long-Horizon Non-Abelian State Tracking

Read the original on arXiv Machine Learning →

arXiv:2606. 07254v1 Announce Type: new Abstract: State tracking exposes a sharp limitation of sequence models: the relevant signal is often not a summary of observed tokens, but an ordered latent state that evolves through non-commutative transformations.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Machine Learning
Aug 17

The Expressive Limits of Diagonal SSMs for State-Tracking

arXiv:2603. 01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable.

By Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath Chandar
arXiv AI
Sep 25

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

The paper investigates how neural networks, particularly Transformers and recurrent models, learn to track group elements by predicting their running product. It finds that Transformers tend to recover quotient classes and that their accuracy can be predicted by the reciprocal of class size, while recurrent networks can capture both normal and non‑normal right‑coset partitions. The study links partial accuracy, learning stages, and internal state representations to the subgroup cosets the models learn to track.

By Zhiyu Zhang, Yupeng Li
arXiv AI
Sep 3

CHASE: Cache-Hole-Adapted Skip Exit for Looped State-Space Language Models

The paper introduces CHASE, a cache‑hole‑adapted skip‑exit mechanism for looped state‑space language models, specifically Looped Mamba and Looped Hybrid Mamba‑Transformer. It shows that looping these architectures improves performance on controlled reasoning tasks and remains competitive in pre‑training benchmarks while using fewer distinct parameters. The cache‑hole adaptation allows selective skipping of recurrent steps during inference, maintaining perplexity close to full computation and achieving significant speedups.

By Zhenxuan Yu, Takeshi Kojima, Yutaka Matsuo, Yusuke Iwasawa