Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.
By Ngoc Bui, Tinglin Huang, Rex Ying
arXiv:2602. 23050v2 Announce Type: replace Abstract: Deep state-space models (DSSMs) enable temporal predictions by learning the underlying dynamics of observed sequence data.
By Alexej Klushyn, Richard Kurle, Maximilian Soelch, Botond Cseke, Patrick van der Smagt
arXiv:2511. 16340v2 Announce Type: replace Abstract: Efficient Gaussian process (GP) inference is critical for sequential decision-making tasks such as active learning, online prediction, and Bayesian optimization.
By Alan Yufei Dong, Jihao Andreas Lin, Jos\'e Miguel Hern\'andez-Lobato
arXiv:2607. 20521v1 Announce Type: new Abstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior.
By Lei Cao, Sihang Feng, Jixin Yan, Tao Sun, Naichen Shi
arXiv:2609.25145v1 Announce Type: cross
Abstract: Variational autoencoders (VAEs) offer an efficient approach to amortized Bayesian inference for inverse problems, but posterior accuracy can depend s...
By Abhishek Srivastava, Arijit Hazra, Rajesh Dubbaku
arXiv:2606. 31063v1 Announce Type: cross Abstract: Gaussian process inference is often limited by cubic computational costs, a challenge that becomes more pronounced in spatio-temporal settings where posterior inference is required over dense grids.
By Rui-Yang Zhang, Lachlan Astfalck, Edward Cripps, David Leslie, Henry Moss