arXiv Machine Learning

Diffusion ReRoll: Revisable Denoising for Robotic Sequential Prediction

arXiv:2607. 19919v1 Announce Type: cross Abstract: We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons.

arXiv Machine Learning
Sep 14

Dynin-Robotics: Omnimodal Unified Diffusion Vision-Language-Action Model

Dynin‑Robotics introduces an omnimodal masked‑diffusion backbone, Dynin‑Omni, that jointly represents language, visual observations, goals, and actions as discrete tokens. By conditioning on different spans, the same model learns action prediction, next‑observation prediction, goal‑state prediction, and trajectory‑to‑instruction reconstruction, enabling test‑time scaling through goal prediction and action‑candidate evaluation. The system, pretrained on 1.33 million trajectories from 48 Open X‑Embodiment datasets, achieves competitive performance on LIBERO, zero‑shot LIBERO‑Plus, and a 78.4 % success rate on a Franka Research 3 robot, while a block‑parallel implementation speeds up action decoding by up to 29.2×.

By Hoeun Lee, Jaeik Kim, Jusang Oh, Jinhyeok Kim, Geon Choi, Hyeonggeun Kim, Jaeyoung Do
arXiv Computer Vision
Sep 7

Persistent Robot World Models: Stabilizing Multi-Step Rollouts via Reinforcement Learning

The paper introduces a reinforcement learning post‑training scheme that trains robot world models on their own autoregressive rollouts, using a contrastive RL objective adapted from diffusion models. It also proposes a training protocol that compares multiple variable‑length futures, a multi‑view visual fidelity reward, and demonstrates state‑of‑the‑art rollout fidelity on the DROID dataset, outperforming baselines on LPIPS, SSIM, and human preference tests.

By Jai Bardhan, Patrik Drozdik, Josef Sivic, Vladimir Petrik
arXiv AI
Aug 19

ManiCM: Real-time 3D Diffusion Policy via Consistency Model for Robotic Manipulation

ManiCM is a real‑time 3D diffusion policy for robotic manipulation that uses a consistency constraint to enable one‑step inference. The model conditions on point‑cloud input and directly predicts robot actions through a consistency distillation technique, avoiding the need to predict noise. Evaluated on 31 tasks from Adroit and Metaworld, ManiCM achieves an average ten‑fold speedup over state‑of‑the‑art methods while maintaining competitive success rates.

By Zifeng Gao, Guanxing Lu, Tianxing Chen, Wenxun Dai, Ziwei Wang, Chao Shang, Wenbo Ding, Yansong Tang
arXiv Computer Vision
Sep 18

Learning Foresight without Explicit Trajectories for 3D Diffusion Policies

The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.

By Zhongbo Zhang, Zaibin Zhang, Yifan Wang, Changbo Yan, Lijun Wang, Huchuan Lu
arXiv Computer Vision
Sep 3

Spatially Aware World Action Model via Geometric Latent Diffusion

The paper introduces Spatially Aware World Action Model (SA‑WAM), a diffusion‑based framework that extends existing World Action Models by incorporating depth information alongside RGB to enable 3‑D‑aware action and future‑state prediction. SA‑WAM repurposes a pretrained video diffusion model, using a nonlinear encoding to map unbounded depth into the tokenizer’s bounded domain, thus preserving pretrained visual priors without 3‑D‑specific fine‑tuning. The model achieves state‑of‑the‑art performance on RoboCasa and LIBERO‑Plus benchmarks and demonstrates superior real‑world performance on a UR5 robotic arm in randomized environments, while also providing analysis linking world‑model prediction quality to rollout success.

By Javier Alejandro Lopetegui Gonzalez, Paul Pacaud, Cordelia Schmid
arXiv AI
Sep 25

Rolling-WAM: World Action Models with Rolling Imagination

Rolling-WAM is a new formulation for World Action Models that spreads the joint video-action denoising process across multiple replanning cycles. It keeps a sliding window of video-action chunks at different noise levels, fully denoising the immediate chunk for execution while partially refining future chunks. This approach reduces latency, improves closed-loop responsiveness, and achieves a 4.5× speedup in steady-state replanning compared to standard WAMs while maintaining competitive manipulation performance.

By Yinghua Zhou, Junjie Ye, Yiqi Zhao, Hao Dong, Celina Shiyu Wang, Ruohai Ge, Tingyi Yang, Basile Van Hoorick, Gaurav Sukhatme, Vitor Guizilini, Yue Wang
Hugging Face Trending Papers
Jun 4

Let It Be Simple: One-Step Action Generation for Vision-Language-Action Models

Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that VLA action generation has a different condition-target structure: the policy is conditioned on rich observations, language, and state, but predicts only a compact, low-dimensional action chunk.