arXiv AI By Yongchao Huang

Your Probabilistic JEPA Is Secretly a Hidden Markov Model: A State-Space Interpretation of Joint-Embedding Predictive Learning

Read the original on arXiv AI →

arXiv:2608. 13621v1 Announce Type: new Abstract: A hidden Markov model (HMM) combines three roles: inference of a hidden-state belief from observations, propagation through a Markov transition, and emission back to observation space.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
5d ago

I Act Therefore I Am: When Is JEPA's Action-Conditioning Enough to Learn Causal Mechanisms?

The paper studies when joint-embedding predictive architectures (JEPAs) can recover underlying causal states from high‑dimensional observations. It introduces a latent variable model where observations arise from causal states with action‑conditioned dynamics, and proposes an information‑theoretic objective that maximizes conditional likelihood while preserving state entropy. The authors prove identifiability conditions—particularly sufficient action‑induced variation—and instantiate the objective as an action‑modulated Gaussian additive‑noise model (A‑JEPA), demonstrating theoretical and empirical success in synthetic and visual benchmarks.

By Yuhang Liu, Zhuo Huang, Javen Qinfeng Shi
arXiv Machine Learning
Jul 27

On the Identifiability of Controlled World Models

arXiv:2607. 22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control.

By Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
arXiv AI
Sep 10

Kalman Delta Networks: Uncertainty-aware Associative Memory

Kalman Delta Networks (KDNs) extend linear attention models by treating associative memory as a linear–Gaussian state‑space system, enabling the Kalman filter to optimally estimate both memory state and its uncertainty. Two GPU‑friendly approximations—Diagonal KDN and Isotropic KDN—use mean‑field variational inference or a single scalar uncertainty per head, respectively, to maintain tractable uncertainty recurrences during linear‑attention scans. Experiments on 750 M and 1.3 B‑parameter models show that KDN variants consistently lower perplexity and raise downstream accuracy compared to existing linear‑attention baselines.

By Ngoc Bui, Tinglin Huang, Rex Ying