arXiv Machine Learning

Fast LeWorldModel

arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.

arXiv AI
3d ago

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv AI
Aug 17

Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost

arXiv:2608. 14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal.

By Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang
arXiv Computer Vision
5d ago

WALT: Learning World-Model-Aligned Latent Trajectories for Autonomous Driving

WALT introduces a method to align latent trajectories with pretrained driving world models, creating a compact generative trajectory space that preserves action-relevant semantics without altering the original model. The approach uses a dual-branch autoencoder to map raw waypoints into this latent space and transfers visual world knowledge into trajectory representations. Experiments on NAVSIM benchmarks show modest performance gains and a 30.5% reduction in planner FLOPs, indicating that maintaining world representations while extracting action-relevant information can improve trajectory planning efficiency.

By Mingkai Jia, Jiaxin Guo, Zhijian Shu, Jiawei Xu, Mingxiao Li, Jintao Cheng, Ping Tan, Wei Yin
arXiv AI
Jul 1

Delta-JEPA: Learning Action-Sensitive World Models via Latent Difference Decoding

arXiv:2606. 31232v1 Announce Type: new Abstract: Learning visual world models for planning requires compact latent dynamics that remain sensitive to actions, yet reconstruction-free joint-embedding objectives can collapse to action-insensitive representations.

By Zhenghao Zhang, Yuanxiang Wang, Zhenyu Guan, Yujia Yang, Bingkang Shi, Tianyu Zong, Hongzhu Yi, Guoqing Chao, Xingchen Chen, Tiankun Yang, Chenxi Bao, Tao Yu, Jingjing Zhou, Jungang Xu
arXiv Machine Learning
Sep 25

Aim Short to Reach Far: Your Frozen World Model Can Plan Better Than You Think

The paper demonstrates that planners using frozen visual world models can achieve better control by changing the target used for action scoring. Instead of scoring actions solely by distance to the final goal image, the authors propose Anchored Planning, which retrieves a recorded trajectory segment that matches the current and goal observations and then scores actions toward an intermediate observation shortly after the segment’s start. Experiments on Cube, PushT, Reacher, and TwoRoom show that this intermediate-target approach outperforms the released LeWM planner on all long‑range tasks, while simple final‑goal search fails to achieve the same gains.

By Xvyuan Liu, Jianjie Fang, Chen Gao, Yong Li
arXiv AI
1d ago

DeepJEPA: Scaling World Models from Within

DeepJEPA is a weight‑tied joint‑embedding predictive world model that treats transition depth as an inner test‑time scaling axis, learning when additional recurrent updates are worthwhile for each candidate and rollout step. Unlike traditional planners that uniformly deepen every transition, DeepJEPA concentrates extra computation on decision‑critical events such as contact onset and sustained object interaction, achieving comparable or better performance with only 1.00–1.26 updates per transition across five visual‑control settings. The approach demonstrates that improved planning does not require uniformly better object‑state decodability, but rather targeted internal computation where it can alter the planner’s elite set and action selection.

By Zijian Jin, Yunbei Zhang, Yuanzhe Liu, Ming Liu, Baian Chen, Weirui Ye, Shilong Liu, Marco Pavone
arXiv AI
Jun 16

LaWAM: Latent World Action Models for Efficient Dynamics-Aware Robot Policies

arXiv:2606. 15768v1 Announce Type: cross Abstract: Vision-Language-Action models (VLAs) leverage large-scale vision-language pretraining for semantic robot control, but often lack explicit foresight into how robot actions change the scene.

By Jialei Chen, Kai Wang, Kang Chen, Shuaihang Chen, Feng Gao, Wenhao Tang, Zhiyuan Li, Weilin Liu, Zhuyu Yao, Boxun Li, Yuanbo Xu, Chao Yu