arXiv AI By Divyansh Jha, Yuanfang Xie, Brennen Yu, Varan Mehra

BRo-JEPA: Learning Modular Transformations in Latent Space

Read the original on arXiv AI →

The paper introduces BRo-JEPA, a world model that learns modular arithmetic operations as rotations in latent space. Using MNIST and EMNIST datasets, BRo-JEPA achieves near-perfect zero‑shot generalization to unseen operations, outperforming standard supervised and JEPA baselines by a large margin. The model demonstrates that latent transformations can encode underlying algebraic structures, enabling strict zero‑shot operation generalization.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 17

Subspace-Decomposed JEPAs: Disentangling Progression and Content in Latent World Models

Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.

By Lucas Thil, Jesse Read, Rim Kaddah, Guillaume Doquet
arXiv AI
Sep 10

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

WorldAgen is a unified framework that jointly learns world modeling and action prediction using a shared Transformer backbone with two specialized heads. It introduces a Mixed Unidirectional Attention Mask to separate the world model and agent model, and enables Test-Time Training (TTT) by sampling exploratory actions and updating the world model with real state transitions. Experiments on CALVIN and LIBERO show that WorldAgen matches or surpasses state‑of‑the‑art methods, especially when TTT is applied to a few samples.

By Chi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang, Manling Li