BRo-JEPA: Learning Modular Arithmetic in Latent Space
arXiv:2606. 01372v1 Announce Type: cross Abstract: Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns?
The paper introduces BRo-JEPA, a world model that learns modular arithmetic operations as rotations in latent space. Using MNIST and EMNIST datasets, BRo-JEPA achieves near-perfect zero‑shot generalization to unseen operations, outperforming standard supervised and JEPA baselines by a large margin. The model demonstrates that latent transformations can encode underlying algebraic structures, enabling strict zero‑shot operation generalization.
arXiv:2606. 01372v1 Announce Type: cross Abstract: Can neural networks learn abstract algebraic rules, or do they merely memorize training patterns?
arXiv:2609.10464v1 Announce Type: cross Abstract: Joint-Embedding Predictive Architecture (JEPA) world models learn a compact latent representation of the world that supports prediction and planning,...
Subspace-Decomposed JEPAs (SD-JEPA) split the latent space of Joint-Embedding Predictive Architectures into two orthogonal subspaces: a low-dimensional progression subspace trained with a cosine-margin triplet loss and a high-dimensional content subspace regularised by SIGReg. The authors prove that the anti-collapse forces act on disjoint coordinates, allowing additive composition rather than competition. SD-JEPA outperforms the LeWM baseline on most control benchmarks and the strongest non-LeWM JEPA baseline on Push‑T, with a subspace-ablation confirming the split as essential. The 1‑D angular progression coordinate serves as a scene-aware compass, advancing with task progress, regressing on backtracking, and relocalising under perturbations to separate surprise from meaning.
WorldAgen is a unified framework that jointly learns world modeling and action prediction using a shared Transformer backbone with two specialized heads. It introduces a Mixed Unidirectional Attention Mask to separate the world model and agent model, and enables Test-Time Training (TTT) by sampling exploratory actions and updating the world model with real state transitions. Experiments on CALVIN and LIBERO show that WorldAgen matches or surpasses state‑of‑the‑art methods, especially when TTT is applied to a few samples.
arXiv:2606. 09936v1 Announce Type: cross Abstract: World models are now built on substantially different computational substrates.
arXiv:2605. 18324v2 Announce Type: replace-cross Abstract: Representation Autoencoders (RAE) replace traditional VAE with pretrained vision encoders.
arXiv:2602. 23164v2 Announce Type: replace Abstract: Foundation models must handle multiple generative processes, yet mechanistic interpretability largely studies capabilities in isolation; it remains unclear how a single transformer organizes multiple, potentially conflicting "world models".
arXiv:2511. 17388v3 Announce Type: replace-cross Abstract: Position information is essential for language modeling.
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2602. 22600v2 Announce Type: replace-cross Abstract: Training selects for behavior, not circuitry: many weight configurations can implement the same function.
arXiv:2609.37250v1 Announce Type: cross Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
arXiv:2607. 06634v1 Announce Type: new Abstract: Compact networks built from Clifford algebra Cl(3,0) primitives are exactly SO(3)-equivariant and learn synthetic 3D vector laws from few samples.