Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.
By Ngoc Bui, Tinglin Huang, Rex Ying
arXiv:2602. 23050v2 Announce Type: replace Abstract: Deep state-space models (DSSMs) enable temporal predictions by learning the underlying dynamics of observed sequence data.
By Alexej Klushyn, Richard Kurle, Maximilian Soelch, Botond Cseke, Patrick van der Smagt
arXiv:2511. 16340v2 Announce Type: replace Abstract: Efficient Gaussian process (GP) inference is critical for sequential decision-making tasks such as active learning, online prediction, and Bayesian optimization.
By Alan Yufei Dong, Jihao Andreas Lin, Jos\'e Miguel Hern\'andez-Lobato
arXiv:2607. 20521v1 Announce Type: new Abstract: The state of a dynamic system evolves over time, switching among several latent modes that govern its observable behavior.
By Lei Cao, Sihang Feng, Jixin Yan, Tao Sun, Naichen Shi
arXiv:2609.25145v1 Announce Type: cross
Abstract: Variational autoencoders (VAEs) offer an efficient approach to amortized Bayesian inference for inverse problems, but posterior accuracy can depend s...
By Abhishek Srivastava, Arijit Hazra, Rajesh Dubbaku
arXiv:2606. 31063v1 Announce Type: cross Abstract: Gaussian process inference is often limited by cubic computational costs, a challenge that becomes more pronounced in spatio-temporal settings where posterior inference is required over dense grids.
By Rui-Yang Zhang, Lachlan Astfalck, Edward Cripps, David Leslie, Henry Moss
The paper introduces a variational framework called VAMO that incorporates latent Markov dynamics for neural PDE solvers, aiming to improve long‑horizon predictions by mitigating error accumulation. By representing physical states as latent distributions and evolving them through probabilistic transitions, the method aligns learned dynamics with a spectral geometry induced by structured Gaussian perturbations. Experiments on fluid‑dynamics benchmarks show that VAMO reduces error growth and enhances rollout stability compared to deterministic and noise‑injection baselines.
By Junyi Liao, Johann Guilleminot, Vahid Tarokh
arXiv:2607. 17614v1 Announce Type: cross Abstract: Recent advances in deep-learning-based nonlinear system identification have led to encoder-based estimation of neural state-space (ANN-SS) models that achieve state-of-the-art performance in offline settings by estimating initial model states from past input-output data.
By Bendeg\'uz Gy\"or\"ok, Tam\'as P\'eni, Maarten Schoukens, Roland T\'oth
arXiv:2601. 07013v2 Announce Type: replace-cross Abstract: Traditional filtering algorithms for state estimation -- such as classical Kalman filtering, unscented Kalman filtering, and particle filters -- show performance degradation when applied to nonlinear systems whose uncertainty follows arbitrary non-Gaussian, and potentially multi-modal distributions.
By Luke S. Lagunowich, Guoxiang Grayson Tong, Daniele E. Schiavazzi
arXiv:2604.07169v3 Announce Type: replace-cross
Abstract: Bayesian filtering and smoothing are central to data assimilation in nonlinear dynamical systems. Recent advances in deep generative models p...
By Tiangang Cui, Xiaodong Feng, Chenlong Pei, Xiaoliang Wan, Tao Zhou
PR‑Smoother is an amortized smoothing method that preserves the explicit use of a prescribed simulator in both the evidence lower bound and the variational family. It learns only future‑conditioned corrections to the simulator’s rollout, yielding a non‑Gaussian smoothing distribution that can jointly infer state, parameters, and sensor bias from observations alone. The approach recovers the exact smoother in deterministic and linear‑Gaussian limits and has been shown to capture multimodal posteriors in Lorenz‑96 and scale to 16,384‑dimensional Kolmogorov flow.
By Yuta Tarumi
arXiv:2606. 01468v1 Announce Type: cross Abstract: Due to their explicit priors and ability to model uncertainty, Bayesian methods have played a major role in dynamical latent variable modeling of single-cell neural recordings.
By JR Huml, Jonathan Wenger, John P. Cunningham