arXiv Machine Learning

Control-Geometry Straightening for Sampling-Based Latent Planning

arXiv Computer Vision
Sep 25

Representation World Model: Learning States, Transition and Executable Plans in Representation

The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.

By Yijun Yuan, Weicheng Zheng, Weibang Wang, Minghui Qin, Chang Sun, Junhao Huang, Kenan Li, Anmin Liu, Yicheng Yao, Hang Zhao
arXiv AI
Aug 17

Traj-LeWM: Path-Aware World-Model Planning via Latent Trajectory Cost

arXiv:2608. 14125v1 Announce Type: new Abstract: LeWM is a lightweight visual world model that learns latent dynamics end-to-end from pixels and ranks candidate action sequences by the distance between their predicted endpoints and the goal.

By Xiaodi Huang, Ziyi Ding, Jingtian Wan, Yuchen Liu, Yuan Zhang, Xiao-Ping Zhang, Jiayu Chen, Zhang Zhang, Tao Huang
arXiv AI
Sep 30

Beyond a single latent space: a dual-latent world model for long-horizon planning

The paper introduces the Dual-Latent World Model (Dual-WM), which separates local execution and long-range planning into distinct latent spaces and dynamics models. A new learning method, Long-Horizon Representation Learning with Weighted Rollout (LoRe), supervises predictions at both levels using exponential horizon weights. Experiments on five goal-conditioned visual control tasks show that Dual-WM improves success rates over strong baselines, especially at longer horizons.

By Delin Zhao, Zhengrong Yue, Shaobin Zhuang, Junlin He, Xiaoyu Chen, Zikang Wang, Yuxin Liu, Limin Wang, Yali Wang
arXiv Machine Learning
2d ago

Reperesentation Geometry Matters for Planning with JEPA World Models

The paper introduces SCALE (State-CAlibrated Latent Embeddings), a technique that aligns pairwise latent distances with task-relevant state-space distances in joint-embedding predictive world models. By adding SCALE to the LeWorldModel (LeWM) objective, the authors preserve LeWM’s architecture while ensuring that latent geometry reflects meaningful task outcomes. Experiments demonstrate that SCALE improves planning success across manipulation and navigation tasks with various solvers, and the authors analyze how the method reshapes representation geometry to support planning.

By Jiaming Hu, Yan Zheng, Shi Bo, Tian Wang, Florian Dubost, Alejandro Mottini, Junze Liu, Arvind Srinivasan, Kai Zhong, Kun Qian, Sharon Gao, Qingjun Cui