arXiv AI By Raktim Gautam Goswami, Prashanth Krishnamurthy, Yann LeCun, Farshad Khorrami

Unifying Object-Centric World Models and Diffusion Policy: A Hierarchical Framework for Multi-Stage Robotic Tasks

Read the original on arXiv AI →

arXiv:2606. 08775v1 Announce Type: cross Abstract: Visual world models have shown great potential in learning complex system dynamics.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Sep 25

Representation World Model: Learning States, Transition and Executable Plans in Representation

The Representation World Model (RWM) learns states, transitions, and executable plans directly within a representation space, bypassing traditional explicit dynamics models and action-space search. It uses inverse-dynamics supervision along latent paths to shape the representation geometry, enabling direct planning by constructing a latent path between current and goal states and recovering actions via inverse dynamics. Experiments on continuous-control benchmarks and robotic manipulation tasks demonstrate RWM’s effectiveness and potential for complex embodied control.

By Yijun Yuan, Weicheng Zheng, Weibang Wang, Minghui Qin, Chang Sun, Junhao Huang, Kenan Li, Anmin Liu, Yicheng Yao, Hang Zhao
arXiv AI
Sep 25

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen is a framework that learns canonical 3D keypoints from point clouds to create structured, object‑centric representations for policy learning in robotic manipulation. By conditioning a visuomotor diffusion policy on these keypoints and object geometry, it predicts full manipulation trajectories that maintain geometric correspondence across different object instances. Experiments on a photorealistic simulation benchmark with three tasks show that KeyGen outperforms prior methods on both seen and unseen objects, scales with more demonstrations, remains robust to rescaling, and performs well in real‑world manipulation.

By Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg
arXiv AI
Jun 3

Coupled Local and Global World Models for Efficient First Order RL

arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.

By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti