arXiv AI

BRo-JEPA: Learning Modular Arithmetic in Latent Space

arXiv:2606. 01372v1 Announce Type: cross Abstract: Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns?

arXiv Machine Learning
Jun 11

Composing Linear Layers from Irreducibles

arXiv:2507. 11688v4 Announce Type: replace Abstract: Contemporary large models often exhibit behaviors suggesting the presence of low-level primitives that compose into modules with richer functionality, but these fundamental building blocks remain poorly understood.

By Travis Pence, Daisuke Yamada, Vikas Singh
arXiv AI
Jun 29

The Context-Ready Transformer

arXiv:2606. 27538v1 Announce Type: cross Abstract: We introduce the context-ready transformer, a new recurrent neural network architecture built from a D-layer transformer block that pre-contextualizes each token before it enters the block.

By Mahesh Godavarti
arXiv Machine Learning
Jun 25

RotRNN: Modelling Long Sequences with Rotations

arXiv:2407. 07239v3 Announce Type: replace Abstract: Linear recurrent neural networks, such as State Space Models (SSMs) and Linear Recurrent Units (LRUs), have recently shown state-of-the-art performance on long sequence modelling benchmarks.

By Kai Biegun, Rares Dolga, Jake Cunningham, David Barber