NaP-Control: Navigating Diffusion Prior for Versatile and Fast Character Control
arXiv:2605. 20209v2 Announce Type: replace-cross Abstract: Achieving precise, versatile whole-body character control in physics-based animation remains challenging.
The paper introduces Action Diffusion, a conditional diffusion model that generates action sequences for a point-mass system affected by dry friction and stiction. Using a compact 1D U‑Net, the model produces bounded control sequences conditioned on initial and target states, outperforming uniform random shooting, structured random shooting, and the Cross‑Entropy Method in reducing terminal error and stuck steps, especially with few samples. The results demonstrate that conditional diffusion can produce temporally coherent controls that effectively overcome stiction by recombining structured primitives from the training prior.
arXiv:2605. 20209v2 Announce Type: replace-cross Abstract: Achieving precise, versatile whole-body character control in physics-based animation remains challenging.
arXiv:2602. 04037v3 Announce Type: replace Abstract: Learning domain adaptive policies that can generalize to unseen transition dynamics, remains a fundamental challenge in learning-based control.
arXiv:2606. 11019v1 Announce Type: cross Abstract: Learning-based motion planners, despite recent progress, often suffer from temporal inconsistency.
The paper introduces Movement Trend Guidance, a method that equips 3D diffusion policies with foresight by learning a compact latent representation of interaction evolution from a brief observation history. This latent, supervised by sparse future gripper states during training, serves as future-oriented conditioning during inference, enhancing action generation without adding explicit planning. The approach improves performance on RoboTwin2.0, LIBERO-40, and DexArt benchmarks, achieving higher success rates across multiple tasks.
Learning-based motion planners, despite recent progress, often suffer from temporal inconsistency. Small perturbations across frames can accumulate into unstable trajectories, degrading comfort and safety in closed-loop driving.
arXiv:2506. 20668v3 Announce Type: replace-cross Abstract: We propose DemoDiffusion, a simple method for enabling robots to perform manipulation tasks by imitating a single human demonstration, without requiring task-specific training or paired human-robot data.
The paper introduces Diffusion-Augmented Markov Decision Processes (DA‑MDPs), a framework that extends Maximum Entropy Reinforcement Learning to diffusion-based policies. DA‑MDPs treat each reverse‑diffusion step as an RL decision, deriving a tractable reverse‑KL bound that decomposes across denoising transitions and yields diffusion‑augmented soft rewards, value functions, and policy objectives. The authors implement this framework with PPO, REPPO, and a maximum‑entropy WPO variant, showing improved continuous‑control performance, higher success rates on manipulation tasks, and memory‑efficient training with action chunking.
arXiv:2607. 19919v1 Announce Type: cross Abstract: We propose Diffusion ReRoll, a diffusion-based framework for robotic sequential prediction that enables revisable denoising over horizons.
arXiv:2606. 08657v1 Announce Type: cross Abstract: Diffusion-based visuomotor policies operating directly in raw action spaces conflate scene comprehension with trajectory generation within a single denoising process.
arXiv:2606. 05737v1 Announce Type: cross Abstract: Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising.
arXiv:2602. 07339v2 Announce Type: replace Abstract: Diffusion-based trajectory planners can model multi-modal driving behavior, but their iterative denoising process introduces a latency bottleneck for real-time closed-loop deployment.
Diffusion-based vision-language-action (VLA) models often inherit the image-generation view: actions are generated by iterative denoising. We argue that VLA action generation has a different condition-target structure: the policy is conditioned on rich observations, language, and state, but predicts only a compact, low-dimensional action chunk.