Hierarchical Planning with Latent World Models
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
arXiv:2604. 03208v2 Announce Type: replace Abstract: World models are a promising path to zero-shot embodied control through planning.
arXiv:2607. 17973v1 Announce Type: new Abstract: Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences.
arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.
arXiv:2608. 14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal.
FlexiWorld is a JEPA-based latent world model that learns variable‑length action chunks across multiple time scales for goal‑directed planning. It jointly trains a causal action encoder and an autoregressive actor, using mixed‑span goal supervision and Student Forcing to reduce exposure bias. In experiments on four benchmarks, FlexiWorld with the Actor‑Residual Cross‑Entropy Method (ARCEM) achieves higher mean success rates than the strongest baseline and supports flexible planning chunk lengths without retraining.
The paper demonstrates that planners using frozen visual world models can achieve better control by changing the target used for action scoring. Instead of scoring actions solely by distance to the final goal image, the authors propose Anchored Planning, which retrieves a recorded trajectory segment that matches the current and goal observations and then scores actions toward an intermediate observation shortly after the segment’s start. Experiments on Cube, PushT, Reacher, and TwoRoom show that this intermediate-target approach outperforms the released LeWM planner on all long‑range tasks, while simple final‑goal search fails to achieve the same gains.
arXiv:2608. 16287v1 Announce Type: new Abstract: Joint-embedding predictive world models plan by scoring predicted terminal embeddings against a goal embedding using a cost defined on the representation itself.
arXiv:2606.27504v2 Announce Type: replace Abstract: World Action Models (WAMs) unify future environment prediction with action generation for autonomous driving, yet existing approaches optimize only...
FIRM-WM is a compact pixel world model that separates a goal‑comparable configuration from a 128‑dimensional dynamic fiber, enabling reward‑free visual planning from offline videos. It addresses two key mismatches: aligning planning states with goal images and reconciling factual trajectories with interventional sampling. In experiments, FIRM‑WM achieves high success rates on TwoRoom, Reacher, and OGBench‑Cube while using fewer parameters and faster planning times than prior models.
arXiv:2605. 08732v2 Announce Type: replace-cross Abstract: Modern vision-based world models can represent observations as compact yet expressive latent manifolds, but fast goal-oriented planning in these spaces remains challenging.
arXiv:2607. 12547v1 Announce Type: cross Abstract: We investigate whether temporal hierarchy can improve LeWorldModel on long-horizon goal-conditioned control.
arXiv:2609.13845v1 Announce Type: cross Abstract: World models trained with joint-embedding predictive architectures learn compact, structured latent representations from physical interaction, yet pl...