arXiv AI

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

arXiv:2608. 13621v1 Announce Type: new Abstract: A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagation through a Markov transition, and emission back to observation space.

arXiv Machine Learning
5d ago

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.

By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv Machine Learning
Jul 27

On the Identifiability of Controlled World Models

arXiv:2607. 22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control.

By Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
arXiv AI
Sep 10

Kalman Delta Networks: Uncertainty-aware Associative Memory

Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.

By Ngoc Bui, Tinglin Huang, Rex Ying
arXiv Machine Learning
Jun 16

Graphical conditional generative modeling for digital twin modeling

arXiv:2606. 16219v1 Announce Type: cross Abstract: Digital twin modeling, including control and data assimilation under model uncertainty, often faces an open-ended fidelity problem: adding variables, data streams, and time scales can indefinitely increase model complexity, ultimately producing systems that are difficult to maintain, validate, interpret, and use for stress or safety testing.

By Zongren Zou, Th\'eo Bourdais, Ricardo Baptista, Houman Owhadi