arXiv Machine Learning By Julien Siems, Riccardo Grazzi, Korbinian P\"oppel, Jaisidh Singh, Arber Zela, Timur Carstensen, Jenia Jitsev, Frank Hutter, Volkan Cevher, Antonio Orvieto, Aaron Klein

Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention

Read the original on arXiv Machine Learning →

The Flow has not summarised this story yet — read it at arXiv Machine Learning.

arXiv Machine Learning
Jun 11

Kalman Linear Attention: Parallel Bayesian Filtering For Efficient Language Modelling and State Tracking

arXiv:2602. 10743v2 Announce Type: replace Abstract: State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking.

By Vaisakh Shaj, Cameron Barker, Aidan Scannell, Andras Szecsenyi, Elliot J. Crowley, Amos Storkey
arXiv AI
Sep 10

Kalman Delta Networks: Uncertainty-aware Associative Memory

Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.

By Ngoc Bui, Tinglin Huang, Rex Ying