The paper introduces AURA, a meta‑learning framework that learns a low‑dimensional latent state‑space model for the evolution of optimal model parameters under distribution shift. Online adaptation is performed via extended Kalman filtering in this latent space, followed by reconstruction of full model parameters through a learned lifting map, enabling efficient single‑step updates. Experiments on neural wireless receivers and non‑stationary image classification show that AURA improves adaptation speed, accuracy, and computational efficiency compared to existing online learning and Bayesian filtering baselines.
By Guy Gerson, Tomer Raviv, Nir Shlezinger, Tirza Routtenberg, Osvaldo Simeone
The paper proposes treating a neural network’s layers as time steps in a state‑space model, converting Bayesian training into a smoothing problem. By propagating Gaussian moments forward and applying a Rauch–Tung–Striebel backward pass, weight posteriors are updated in closed form without gradient iterations or replay. The authors extend prior work by introducing a cross‑covariance identity that allows full‑covariance propagation through nonlinear activations, enabling more accurate online adaptation in non‑stationary classification, dynamics learning, and vision‑language‑action policy adaptation.
By Oren Wright, Haoming Jing, Qiaoan Shen, Koichiro Niinuma, Yorie Nakahira, Jos\'e M. F. Moura
The paper studies a two‑stage learning framework that first trains an offline model using approximate nonlinear‑least‑squares estimation and then adapts it online with a meta‑LMS algorithm to handle parameter drift in nonlinear stochastic dynamical systems. It provides an upper bound on the offline generalization error that accounts for strong data correlation and distribution shift via Kullback‑Leibler divergence, and it demonstrates that the combined offline‑online approach outperforms methods that rely solely on offline or online learning. Both theoretical analysis and empirical experiments support the claimed performance gains.
By Haizheng Li, Lei Guo
arXiv:2602. 23050v2 Announce Type: replace Abstract: Deep state-space models (DSSMs) enable temporal predictions by learning the underlying dynamics of observed sequence data.
By Alexej Klushyn, Richard Kurle, Maximilian Soelch, Botond Cseke, Patrick van der Smagt
arXiv:2609.36712v1 Announce Type: cross
Abstract: Accurately learning nonlinear dynamics from a finite-duration experiment requires the efficient collection of informative data. We address this chall...
By Juncal Arbelaiz, Anushri Arora, Jonathan W. Pillow
arXiv:2607. 06079v1 Announce Type: new Abstract: Intelligent systems should not only solve tasks but also adapt under real-world constraints.
By Taiki Yamada, Kantaro Fujiwara
arXiv:2412. 12036v2 Announce Type: replace Abstract: System identification, the process of deriving mathematical models of dynamical systems from observed input-output data, has undergone a paradigm shift with the advent of learning-based methods.
By Arunabh Singh, Joyjit Mukherjee
arXiv:2503. 18970v4 Announce Type: replace Abstract: Structured State Space Models (SSMs) have become a prominent class of sequence models, developed against two long-standing difficulties: the sequential computation and gradient propagation limits of Recurrent Neural Networks (RNNs), and the quadratic time and memory cost of self-attention in Transformers.
By Shriyank Somvanshi, Md Monzurul Islam, Mahmuda Sultana Mimi, Sazzad Bin Bashar Polock, Gaurab Chhetri, Anandi Dutta, Amir Rafe, Subasish Das
arXiv:2511. 15409v2 Announce Type: replace Abstract: We present a class of algorithms for state estimation in nonlinear, non-Gaussian state-space models.
By Hany Abdulsamad, \'Angel F. Garc\'ia-Fern\'andez, Simo S\"arkk\"a
RD‑JEPA is a joint‑embedding predictive architecture designed for self‑supervised pretraining on reaction‑diffusion trajectories. The model is pretrained on five parameterized systems and then adapted to three held‑out systems that were not seen during pretraining. Using as few as one, five, or ten complete trajectories from a held‑out system, RD‑JEPA outperforms five supervised surrogate baselines, an independently trained control that removes the trajectory‑dependent predictive latent pathway, and an architecture‑matched model trained from scratch, achieving lower mean relative discrete β field error and mean absolute spatial first‑difference error across various output resolutions, forecast horizons, and adaptation trajectory choices.
By Chenhao Si, Ming Yan
The paper introduces Linearized Subspace Refinement (LSR), a post‑training framework that uses the local linearized model of a trained neural network to compute a low‑dimensional correction via a reduced least‑squares problem. LSR is architecture‑agnostic and improves accuracy across tasks such as function approximation, operator learning, physics‑informed fine‑tuning, and noisy inverse problems, often achieving order‑of‑magnitude error reductions. The method reveals that standard training can leave significant accuracy plateaus due to numerical ill‑conditioning, and it offers a subspace rank that balances correction strength, stability, and noise sensitivity.
By Wenbo Cao, Weiwei Zhang
arXiv:2607. 18965v1 Announce Type: cross Abstract: Deep learning has proven highly effective for nonlinear system identification, but heavily parameterized neural networks are prone to overfitting in low-data regimes and lack reliable uncertainty quantification.
By Matteo Rufolo, Dario Piga, Marco Forgione