arXiv AI

ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

ManiCM is a real‑time 3D diffusion policy for robotic manipulation that uses a consistency constraint to enable one‑step inference. The model conditions on point‑cloud input and directly predicts robot actions through a consistency distillation technique, avoiding the need to predict noise. Evaluated on 31 tasks from Adroit and Metaworld, ManiCM achieves an average ten‑fold speedup over state‑of‑the‑art methods while maintaining competitive success rates.

arXiv Machine Learning
Jun 5

Flash-WAM: Modality-Aware Distillation for World Action Models

arXiv:2606. 05254v1 Announce Type: new Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control.

By Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang
arXiv AI
Sep 23

HybridFlow: A 2-NFE Generative Policy for Real-Time Robotic Manipulation

HybridFlow is a generative policy for robotic manipulation that uses a three‑stage inference procedure requiring only two network function evaluations (2‑NFE). The policy first generates a coarse action trajectory with a Global Jump based on MeanFlow, then refines the state using a parameter‑free ReNoise interpolation, and finally performs a Local Refine to query the instantaneous‑velocity limit. Experiments on RoboMimic and five real‑robot settings show that HybridFlow achieves high success rates and improves task performance over a 16‑step Diffusion Policy while reducing action‑generation latency by roughly eightfold.

By Zhenchen Dong, Fulin Chen, Jinna Fu, Jiaming Wu, Qingran Wu, Shengyuan Yu, Hongyu Yu, Yide Liu
arXiv AI
Jun 3

Coupled Local and Global World Models for Efficient First Order RL

arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.

By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti
arXiv AI
Sep 25

Rolling-WAM: World Action Models with Rolling Imagination

Rolling-WAM is a new formulation for World Action Models that spreads the joint video-action denoising process across multiple replanning cycles. It keeps a sliding window of video-action chunks at different noise levels, fully denoising the immediate chunk for execution while partially refining future chunks. This approach reduces latency, improves closed-loop responsiveness, and achieves a 4.5× speedup in steady-state replanning compared to standard WAMs while maintaining competitive manipulation performance.

By Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
arXiv Computer Vision
Sep 3

Spatially Aware World Action Model via Geometric Latent Diffusion

The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.

By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid