Orthogonal JEPA: Factorized Predictive States for Latent World Models
arXiv:2608. 20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system.
Orthogonal JEPA introduces a latent world‑modeling framework that factorizes predictive states into orthogonal components. By learning basis matrices and dedicated prediction branches, the method reduces redundancy and improves gradient signals for less dominant predictive structures. The factorized states can be synthesized into complete latent representations for downstream tasks such as decoding, planning, or autoregressive rollout, and are evaluated across vision, biology, health, control, and molecular dynamics domains.
arXiv:2608. 20065v1 Announce Type: new Abstract: World models construct latent states that support prediction, planning, and reasoning about an underlying system.
JEPA-Anything is a domain‑agnostic framework that uses orthogonal predictive factorization (OPF) to decompose latent targets into complementary factors, learn them via dedicated pathways, and recombine them for shared prediction. The method is evaluated across seven diverse domains—vision, biology, clinical trajectories, control, molecular dynamics, physical fields, and weather—showing improvements on 10 dynamics tasks, reduced error on Interventional Pong, and lowest one‑step and 100‑step molecular errors among compared methods. Experimental validation includes a factor‑nominated biological intervention that succeeded in cell co‑cultures, organoids, tumor fragments, and mice, and latent orbital modes that recover the Keplerian scaling exponent.
Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.
Joint-embedding predictive architectures (JEPAs) learn latent dynamics for planning and avoid representation collapse by matching features to maximum-entropy distributions such as isotropic Gaussians,...
arXiv:2606. 27014v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have recently emerged as a promising paradigm for world modeling by learning predictive dynamics in a latent space rather than generating future observations at the input level.
WorldTS is a new forecasting framework that models latent dynamics conditioned on multimodal covariates to improve time‑series prediction. It uses a two‑stage training process: first learning latent state dynamics from historical data and covariates, then training a decoder to map predicted latent states back to future observations. Experiments on 21 real‑world datasets demonstrate the effectiveness of this approach.
arXiv:2605. 10840v3 Announce Type: replace-cross Abstract: We present Clin-JEPA, a multi-phase co-training framework for joint-embedding predictive (JEPA) pretraining on EHR patient trajectories.
arXiv:2608. 16287v1 Announce Type: new Abstract: Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself.
arXiv:2608.29029v3 Announce Type: replace-cross Abstract: Joint-Embedding Predictive Architectures (JEPAs) provide a powerful framework for latent world modeling and planning in a reconstruction-free...
arXiv:2606. 16076v1 Announce Type: cross Abstract: Multivariate forecasting in physical systems requires models that predict coupled temporal variables while preserving meaningful state evolution.
arXiv:2608. 05928v1 Announce Type: new Abstract: Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes.
Single-cell transcriptomes are sparse observations of coordinated biological programmes, yet most self-supervised models learn by reconstructing individual genes. Here we present BioM-JEPA, a joint-embedding predictive architecture that instead predicts aggregate representations of graph-connected gene blocks defined by protein-association and corpus-derived coexpression evidence.