arXiv AI By Zhiyu Zhang, Yupeng Li

Tracking States or Tracking Cosets? An Algebraic Account of Learned State Tracking

Read the original on arXiv AI →

The paper investigates how neural networks, particularly Transformers and recurrent models, learn to track group elements by predicting their running product. It finds that Transformers tend to recover quotient classes and that their accuracy can be predicted by the reciprocal of class size, while recurrent networks can capture both normal and non‑normal right‑coset partitions. The study links partial accuracy, learning stages, and internal state representations to the subgroup cosets the models learn to track.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Aug 17

The Expressive Limits of Diagonal SSMs for State-Tracking

arXiv:2603. 01959v2 Announce Type: replace Abstract: State-Space Models (SSMs) have recently been shown to achieve strong empirical performance on a variety of long-range sequence modeling tasks while remaining efficient and highly-parallelizable.

By Mehran Shakerinava, Behnoush Khavari, Siamak Ravanbakhsh, Sarath Chandar
arXiv AI
Sep 21

Understanding In-context Learning of Addition via Activation Subspaces

The paper investigates how transformer language models perform few‑shot learning for a simple addition task, showing that the ability is concentrated in a handful of attention heads. Using dimensionality reduction, the authors identify low‑dimensional subspaces—three heads with six‑dimensional spaces in Llama‑3‑8B‑Instruct—where specific dimensions encode the units digit via trigonometric patterns and magnitude via low‑frequency components. They also derive a mathematical identity linking aggregator and extractor subspaces, enabling tracking of information flow from examples to the final prediction.

By Xinyan Hu, Kayo Yin, Michael I. Jordan, Jacob Steinhardt, Lijie Chen