arXiv:2607. 19919v1 Announce Type: cross Abstract: We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons.
By Seonsoo Kim, Seongil Hong, Jun-Gill Kang
arXiv:2606. 08657v1 Announce Type: cross Abstract: Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process.
By Zhexuan Zhou, Yichen Lai, Jinhao Zhang, Huizhe Li, Youmin Gong, Jie Mei
arXiv:2606. 03943v1 Announce Type: cross Abstract: Video-Action Models (VAMs) leverage the broad visual dynamics captured by pre-trained video diffusion models, offering a promising path toward generalizable robot manipulation.
By Mutian Tong, Han Jiang, Qiao Feng, Lingjie Liu, Jiatao Gu
arXiv:2512. 07212v3 Announce Type: replace Abstract: Imitation learning with diffusion models has advanced robotic control by capturing the multi-modal action distributions.
By Zhaoyang Liu, Mokai Pan, Zhongyi Wang, Kaizhen Zhu, Haotao Lu, Haipeng Zhang, Jingya Wang, Ye Shi
arXiv:2606. 05254v1 Announce Type: new Abstract: World-action models (WAMs) jointly generate future video and robot actions through iterative diffusion, achieving strong performance on manipulation benchmarks but requiring tens of denoising steps, a cost that precludes real-time control.
By Arman Akbari, Ci Zhang, Arash Akbari, Lin Zhao, Yixiao Chen, Weiwei Chen, Xuan Zhang, Geng Yuan, Yanzhi Wang
HybridFlow is a generative policy for robotic manipulation that uses a three‑stage inference procedure requiring only two network function evaluations (2‑NFE). The policy first generates a coarse action trajectory with a Global Jump based on MeanFlow, then refines the state using a parameter‑free ReNoise interpolation, and finally performs a Local Refine to query the instantaneous‑velocity limit. Experiments on RoboMimic and five real‑robot settings show that HybridFlow achieves high success rates and improves task performance over a 16‑step Diffusion Policy while reducing action‑generation latency by roughly eightfold.
By Zhenchen Dong, Fulin Chen, Jinna Fu, Jiaming Wu, Qingran Wu, Shengyuan Yu, Hongyu Yu, Yide Liu
arXiv:2602. 06219v2 Announce Type: replace-cross Abstract: World models offer a promising avenue for more faithfully capturing complex dynamics, including contacts and non-rigidity, as well as complex sensory information, such as visual perception, in situations where standard simulators struggle.
By Joseph Amigo, Rooholla Khorrambakht, Nicolas Mansard, Ludovic Righetti
Rolling-WAM is a new formulation for World Action Models that spreads the joint video-action denoising process across multiple replanning cycles. It keeps a sliding window of video-action chunks at different noise levels, fully denoising the immediate chunk for execution while partially refining future chunks. This approach reduces latency, improves closed-loop responsiveness, and achieves a 4.5× speedup in steady-state replanning compared to standard WAMs while maintaining competitive manipulation performance.
By Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.
By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv:2607. 15065v1 Announce Type: cross Abstract: Predictive world models enable robots to plan by imagining the outcomes of their actions, but their value for control hinges on generating many rollouts quickly.
By Susie Lu, Haonan Chen, Weirui Ye, Yilun Du
arXiv:2607. 10706v1 Announce Type: cross Abstract: The action space poses a major challenge in robot learning, since it is often high-dimensional, can span long time horizons, and frequently admits multi-modal optimal solutions.
By Haojie Huang, Zhang Ye, Linfeng Zhao, Boce Hu, Mingxi Jia, Yu Qi, Ahmed Agha, Dian Wang, Robert Platt, Robin Walters
arXiv:2506. 20668v3 Announce Type: replace-cross Abstract: We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or paired human-robot data.
By Sungjae Park, Homanga Bharadhwaj, Shubham Tulsiani