arXiv:2607. 13031v1 Announce Type: new Abstract: When one ball strikes another, then another, video models should predict the consequences of each bounce.
By Jorge Diaz Chao, Konpat Preechakul, Yuxi Liu, Yutong Bai
arXiv:2406.04814v4 Announce Type: replace-cross
Abstract: Video diffusion models can enable embodied agents to anticipate plausible futures from the recent past, but they are typically trained offlin...
By Jason Yoo, Yingchen He, Saeid Naderiparizi, Dylan Green, Gido M. van de Ven, Geoff Pleiss, Frank Wood
S2PD: Serial-to-Parallel Diffusion for Physically and Logically Consistent Video Generation introduces a hybrid diffusion approach that first applies autoregressive diffusion at high noise levels and then switches to parallel diffusion at low noise levels. This method coordinates interdependent events to produce valid state transitions while reducing sampling time compared to fully serial generation. Implemented with a pixel‑space diffusion transformer and a LoRA‑fine‑tuned pretrained video model, S2PD outperforms bidirectional baselines in rule adherence and achieves greater temporal stability and sampling efficiency across games, physical simulations, and real video.
LiveVVT introduces a rolling streaming diffusion framework for video virtual try‑on that maintains high visual fidelity while enabling real‑time performance. It preserves bounded bidirectional modeling within a fixed‑size window, emits clean video chunks iteratively, and uses two memory modules—a bounded temporal memory and a persistent global appearance memory—to sustain long‑term consistency. A progressive distillation process further aligns teacher‑based bidirectional learning with causal few‑step inference, resulting in superior generation quality with 26× lower latency and 11× higher throughput compared to comparable models.
By Yushe Cao, Shikun Feng, Ruxiang Duan, Liyong Wang, Dianxi Shi, Chun Yu, Junliang Xing
arXiv:2609.23658v1 Announce Type: cross
Abstract: Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While exist...
By Yueyan Li, Haibo Wang, Caixia Yuan, Xiaojie Wang
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.