arXiv Computer Vision

R3D: Revisiting 3D Policy Learning

arXiv Computer Vision
Sep 3

Spatially Aware World Action Model via Geometric Latent Diffusion

The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.

By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
Hugging Face Trending Papers
Jun 1

RoboDream: Compositional World Models for Scalable Robot Data Synthesis

Scaling robot learning requires large-scale, diverse demonstrations, yet real-world data collection via teleoperation remains prohibitively expensive and time-consuming. While video diffusion models offer a promising avenue for data scaling, existing generative approaches are often limited to superficial visual augmentation, or suffer from embodiment hallucinations that yield physically infeasible motions.

arXiv AI
Sep 25

KeyGen: Unsupervised Keypoint based Object-Centric Representations for Category-Level Policy Generalization

KeyGen is a framework that learns canonical 3D keypoints from point clouds to create structured, object‑centric representations for policy learning in robotic manipulation. By conditioning a visuomotor diffusion policy on these keypoints and object geometry, it predicts full manipulation trajectories that maintain geometric correspondence across different object instances. Experiments on a photorealistic simulation benchmark with three tasks show that KeyGen outperforms prior methods on both seen and unseen objects, scales with more demonstrations, remains robust to rescaling, and performs well in real‑world manipulation.

By Shuxin Cao, Liquan Wang, Masoud Moghani, Benjamin Joffe, Animesh Garg
arXiv AI
Jun 3

Coupled Local and Global World Models for Efficient First Order RL

arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.

By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti