arXiv:2607. 27036v1 Announce Type: cross Abstract: Video diffusion-based world models enable long autoregressive video generation for robotics, autonomous driving and simulation tasks, yet sliding-window autoregressive inference suffers from severe error accumulation that degrades frame quality over time.
By Taiye Chen, Qi Zhang, Yisen Wang
The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.
By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
arXiv:2606. 14732v1 Announce Type: cross Abstract: Autoregressive video diffusion models enable streaming generation but often degrade over long rollouts: static scene layouts drift, while mechanisms that improve spatial stability tend to suppress motion, causing natural flows such as water, fire, or smoke to stagnate.
By Matiur Rahman Minar, Seunghun Oh, GangHyeon Jeong, Unsang Park
Manifold4D introduces a new denoising strategy for video re‑shooting that injects a rendered point‑cloud directly into the initial noise manifold, eliminating the need for the render to be an explicit conditioning stream during denoising. This approach allows the network to rely solely on the source video as a visual condition, improving camera‑control accuracy on the DAVIS‑Traj benchmark and Vista4D set, with significant reductions in rotation and translation errors while maintaining video fidelity. User studies confirm enhanced trajectory following and dynamic consistency, especially for large yaw amplitudes and even when the render is corrupted.
By Yongqi Mao, Zijia Dai, Zhishuo Liu, Wei Xu, Kaiwei Wang, Guotao Meng
arXiv:2607. 15849v1 Announce Type: cross Abstract: Autoregressive video diffusion models have enabled the generation of arbitrarily long videos by removing conditioning on future frames, thus greatly improving computational efficiency.
By Dimitrios Karageorgiou, Symeon Papadopoulos, Ioannis Kompatsiaris, Efstratios Gavves
arXiv:2603. 14294v3 Announce Type: replace-cross Abstract: Do video diffusion models encode signals predictive of physical plausibility?
By Chujun Tang, Lei Zhong, Fangqiang Ding
arXiv:2608. 19556v1 Announce Type: cross Abstract: Streaming autoregressive diffusion models enable real-time, long-horizon video generation, but their training objectives optimize local frame prediction rather than the geometry and dynamics of a coherent world: long rollouts accumulate geometric drift and degrade into static or unnatural motion.
By Yuanhao Ban, Jiaqi Feng, Hengguang Zhou, Xiaohuan Pei, Justin Cui, Cho-Jui Hsieh
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2608. 05237v1 Announce Type: cross Abstract: Current few-step autoregressive video diffusion models depend on previous fully denoised clean frames as context for all denoising steps of the current frame.
By Lingxiao Yang, Liu Liu, Moran Li, Han Feng, Wenjian Cao, Jiangning Zhang, Ye Shi
arXiv:2609.38114v1 Announce Type: new
Abstract: Autoregressive video diffusion enables interactive streaming generation, but suffers from error accumulation over long rollouts. Self-rollout training...
By Weiqiang Wang, Zhuokun Chen, Yusheng Dai, Boying Li, Yi Zhang, Hossein Rahmani, Qiuhong Ke, Jianfei Cai
MotionSpec introduces a motion supervision framework for text-to-video generation that focuses on Spectral Trajectory Consistency (STC). STC builds dense anchor-relative motion trajectories, transforms them into spectral volumes, and aligns their amplitude and phase with target trajectories to constrain motion strength and temporal organization. The framework also adds Local Flow Consistency (LFC) to stabilize local motion transitions, resulting in improved motion consistency, temporal coherence, and plausibility while maintaining visual fidelity.
By Ziqi Ni, Rui Li, Shiqi Jiang, Wei Zhou
arXiv:2610.00812v1 Announce Type: cross
Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamic...
By Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli