Seeking Physics in Diffusion Noise
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
arXiv:2607. 15849v1 Announce Type: cross Abstract: Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency.
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
arXiv:2609.38114v1 Announce Type: new Abstract: Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training...
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
arXiv:2608.29322v1 Announce Type: new Abstract: Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or c...
arXiv:2607. 27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time.
Manifold4D introduces a new denoising strategy for video re‑shooting that injects a rendered point‑cloud directly into the initial noise manifold, eliminating the need for the render to be an explicit conditioning stream during denoising. This approach allows the network to rely solely on the source video as a visual condition, improving camera‑control accuracy on the DAVIS‑Traj benchmark and Vista4D set, with significant reductions in rotation and translation errors while maintaining video fidelity. User studies confirm enhanced trajectory following and dynamic consistency, especially for large yaw amplitudes and even when the render is corrupted.
Accelerating Video Diffusion via Training-Free Trajectory Routing (TRACK) introduces a heterogeneous denoising strategy that switches between large and small diffusion models at selected steps, determined by a calibration process that measures disagreement between model predictions. By routing quality-sensitive steps to the large model and low-disagreement steps to the small model, TRACK achieves significant speedups—up to 2.73×—across several video diffusion benchmarks while maintaining comparable quality and diversity. The method requires no retraining, architectural changes, or online dual-model evaluation, making it a practical acceleration paradigm for video diffusion.
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.
We propose OPSD-V, an on-policy self-distillation paradigm for post-training few-step autoregressive (AR) video diffusion models. Existing few-step AR video generators can produce long videos with low latency, but still suffer from error accumulation and weakened motion dynamics during long autoregressive rollout.
arXiv:2609.26792v1 Announce Type: cross Abstract: Faithfully evaluating end-to-end driving policies in simulation requires observations that are not merely photo-realistic, but preserve the scene fea...
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.