Complex KDA: Understanding and Enhancing the Expressivity of Kimi Delta Attention
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 11897v1 Announce Type: new Abstract: Linear attention replaces softmax attention's growing KV cache with a fixed recurrent state, but this compression limits exact state tracking and long-context memory.
arXiv:2608.30386v1 Announce Type: cross Abstract: Hybrid linear-attention architectures have recently scaled to large open-weight models, offering quality competitive with full attention while substa...
arXiv:2604.19021v3 Announce Type: replace Abstract: Linear attention mechanisms have emerged as promising alternatives to softmax attention, offering linear-time complexity during inference. Recent a...
arXiv:2608.30310v1 Announce Type: cross Abstract: Hybrid large language models interleave full-attention layers with linear-attention layers to reduce the cost of long-context inference. This structu...
arXiv:2602. 10743v2 Announce Type: replace Abstract: State-space language models such as Mamba and gated linear attention (GLA) offer linear-complexity, parallelisable alternatives to transformers, but their linear state updates limit expressivity and robust state tracking.
Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.