Does Latent Planning Survive Point Clouds? Action-Conditioned JEPA World Models for Geometric Observations
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to improve goal‑conditioned robotic planning. By preventing latent collapse and grounding representations in physical configuration, the model achieves top success rates on tasks such as TwoRoom, PushT, and OGBench‑Cube, outperforming the baseline LeWorldModel. Ablation studies confirm that state alignment consistently boosts planning success over inverse dynamics alone across all four benchmark tasks.
arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
arXiv:2606. 07974v1 Announce Type: cross Abstract: A learned world model provides a powerful physical intuition for evaluating future states.
arXiv:2608.24855v1 Announce Type: new Abstract: Latent world models are inherently strong encoders that transform image pixel to latent embedding, yet existing world models still rely on online traje...
The paper presents an end‑to‑end JEPA world model that enhances latent prediction with inverse dynamics and state alignment to better support goal‑conditioned robotic planning. By incorporating inverse dynamics, the model prevents latent collapse and encodes action information, while state alignment ties consecutive latent states to their physical configurations and motions. Experiments on four benchmark tasks show the model achieves top success rates on TwoRoom, PushT, and OGBench‑Cube, and its state alignment consistently improves planning performance over inverse dynamics alone.