The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.
By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
The paper demonstrates that planners using frozen visual world models can achieve better control by changing the target used for action scoring. Instead of scoring actions solely by distance to the final goal image, the authors propose Anchored Planning, which retrieves a recorded trajectory segment that matches the current and goal observations and then scores actions toward an intermediate observation shortly after the segment’s start. Experiments on Cube, PushT, Reacher, and TwoRoom show that this intermediate-target approach outperforms the released LeWM planner on all long‑range tasks, while simple final‑goal search fails to achieve the same gains.
By Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li
PACT‑WAM is a world‑action model that simultaneously predicts a 16‑step action trajectory and its corresponding visual forecast for robot manipulation. It uses a hierarchical history encoder that compresses past observations into fewer tokens, reducing processing cost by 75% compared to dense encoding. The model’s shared flow module updates action and visual states jointly, and a TiTok‑VAE decoder reconstructs multi‑view future images, which are then used by a vision‑language component (Proposal Review) to improve execution‑prefix selection and proposal rejection, boosting success rates on several benchmarks.
By Yushan Liu, Jingjing Fan, Shoujie Li, Yifan Xie, Xiao-Ping Zhang, Wenbo Ding
arXiv:2606. 09028v1 Announce Type: cross Abstract: Latent world models are increasingly used for control and goal-conditioned planning, yet assessing whether their learned representations are useful for planning usually requires slow, planner-coupled simulator evaluation with CEM or similar planners.
By Jiaheng Chen
arXiv:2608. 12939v1 Announce Type: new Abstract: Joint-embedding predictive architectures (JEPAs) learn world models that predict in a compact latent space rather than in pixels, reducing the pressure to model nuisance appearance.
By Guo An, Zijing Wu, Honghua Dong, Yuhao Yan, Zixuan Gui, Haochong Chen, Shanzhao Ruan, Xiang Wang, Yurong Ling, Qi Tian
arXiv:2606. 03685v1 Announce Type: cross Abstract: Supervised fine-tuning (SFT) improves end-to-end classical planning in large language models (LLMs), but do these models also learn to represent and reason about the planning problems they are solving?
By Patrick Emami, Nan Qiang, Peter Graf
World-Coherent Decoding (WCD) is a test-time planning framework for World Action Models (WAMs) that treats rollouts as falsifiable future–action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using flow-based video surprisal for visual plausibility and action path effort for generation stability. After execution, the observed outcome audits the chosen imagination, producing a mismatch signal that trains a lightweight online predictor to improve future candidate selection, thereby enhancing reliability without updating the backbone model.
By Chuhan Zhang, Seiji Ito, Kenta Hoshino, Satoshi Ikehata, Ikuro Sato
arXiv:2609.05834v1 Announce Type: new
Abstract: World models promise a general route to embodied intelligence: learn predictive dynamics once, then reason, plan, and act with them. Increasingly, the...
By Todd Y. Zhou, Daniel Zhang
arXiv:2607. 17973v1 Announce Type: new Abstract: Latent world models have emerged as a powerful planning paradigm by learning action-conditioned predictive dynamics and using them as internal simulators to imagine and evaluate candidate action sequences.
By Letian Cheng, Qi Zhang, Yisen Wang
arXiv:2606. 27326v1 Announce Type: new Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics.
By Nicklas Hansen, Xiaolong Wang
arXiv:2607. 02403v1 Announce Type: cross Abstract: Decision-time planning with action-conditioned world models has become a popular paradigm for embodied control.
By Gawon Seo, Dongwon Kim, Suha Kwak
arXiv:2606. 13053v1 Announce Type: cross Abstract: Pretrained-feature world models provide a useful substrate for robot imagination, but visual or latent prediction alone does not determine whether an imagined future satisfies task-relevant events.
By Kailin Wang, Haoxiang Jie, Yaoyuan Yan, Jiacheng Zhou, Zhiyou Heng