arXiv:2606. 25473v1 Announce Type: cross Abstract: Autoregressive video diffusion with causal diffusion transformers has emerged as a major paradigm for real-time streaming video generation and action-conditioned interactive world models.
By Kaiwen Zheng, Guande He, Min Zhao, Jintao Zhang, Huayu Chen, Jianfei Chen, Chen-Hsuan Lin, Ming-Yu Liu, Jun Zhu, Qianli Ma
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both...
Real-time long-form avatar audio--video generation requires causal, continuous synthesis while maintaining audiovisual synchronization and visual consistency. Adapting a pretrained bidirectional model to this setting presents two key dilemmas.
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
By Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang
arXiv:2609.35491v2 Announce Type: replace-cross
Abstract: Few-step autoregressive video generation commonly relies on Distribution Matching Distillation (DMD), requiring a bidirectional diffusion tea...
By Chi Zhang, Yueyi Liu, Haoyang Shi, Ruichuan An, Haoyu Li, Yuhang Wu, Sen Cui, Miao Liu
arXiv:2609.34581v2 Announce Type: replace
Abstract: Temporal video grounding is a key capability of advanced Multimodal Large Language Models (MLLMs) for the thorough understanding of video events, w...
By Shaobo Ju, Haiyang Yu, Xuecheng Wu, Qiong Wu, Jiacong Wang, Fan Shi, Jun Peng, Yiyi Zhou
arXiv:2603. 01006v3 Announce Type: replace-cross Abstract: REPresentation Alignment (REPA) improves the training of generative flow models by aligning intermediate hidden states with pretrained teacher features, but its effectiveness in token-conditioned audio Flow Matching critically depends on the choice of supervised layers, which is typically made heuristically based on the depth.
By Pengfei Zhang, Tianxin Xie, Minghao Yang, Li Liu
Autoregressive video diffusion models have emerged as a promising approach for long video generation, achieving strong performance in streaming settings. However, existing methods are restricted to forward temporal generation, whereas practical video creation often requires flexible generation order, e.
arXiv:2609.22283v1 Announce Type: new
Abstract: Understanding the design space of streaming video diffusion is essential to exploring its potential for generation quality and computational efficiency...
By Hongchen Zhang (University of Chinese Academy of Sciences)
The paper introduces Prediction‑Aligned Context Compaction (PACC), a method that learns a compact memory representation for long‑video generation by distilling a frozen video generator. PACC trains a compressor to aggregate past frames into memory tokens, using the generator as both teacher and student during on‑policy distillation. Experiments on MBench and VBench‑Long show that PACC improves memory‑event coverage and consistency, achieving better scores than strong baselines and producing competitive minute‑long videos.
By Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng, Lydia E. Kavraki, Zhiwei Steven Wu
arXiv:2605. 30116v2 Announce Type: replace-cross Abstract: Distribution Matching Distillation (DMD) is a widely used paradigm for accelerating inference in few-step video diffusion models.
By Zhuguanyu Wu, Ruihao Gong, Yang Yong, Yushi Huang, Xiangyu Fan, Lei Yang, Dahua Lin, Xianglong Liu
arXiv:2609.23010v1 Announce Type: new
Abstract: Iterative text-to-motion generation delivers high-quality and semantically aligned motions but requires multiple network evaluations, resulting in subs...
By Hung Dinh, Binh Mai, Tran Quoc Bao Le, Lam Nguyen, Cong Tran