arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.
By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.
arXiv:2602.10639v2 Announce Type: replace
Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what m...
By Yuxin Cao, Wei Song, Shangzhi Xu, Jingling Xue, Jin Song Dong
arXiv:2607. 14194v1 Announce Type: cross Abstract: Text-to-video (T2V) generators can synthesize realistic and temporally coherent videos, but controllably removing a target concept from a generator remains difficult.
By Wenxuan Chen, Wenjie Feng
The paper introduces RAE-CoD, a diffusion-based compression method that operates in a representation autoencoder space to preserve recognizable content even at extremely low bitrates. It addresses the problem of semantic collapse observed in existing codecs when the bitrate approaches zero, showing that reconstruction losses conflict with semantic objectives and that VAE diffusion models lose efficiency in preserving semantics. Experiments on MSCOCO-30K demonstrate that RAE-CoD outperforms competitors, reducing VFM feature MSE and Fréchet Distance ratios by at least 25.7% and 69.1% at 0.001–0.008 bpp while maintaining stable recognizability and quality.
By Tianyu Zhang, Zhaoyang Jia, Houqiang Li, Dong Liu
arXiv:2608.29322v1 Announce Type: new
Abstract: Recent video diffusion models have achieved remarkable generation quality, but high-fidelity results still largely depend on closed-source systems or c...
By Hangzhou He, Lunhao Duan, Shanshan Zhao, Kaiwen Li, Qing-Guo Chen, Weihua Luo, Yanye Lu
arXiv:2610.00686v1 Announce Type: new
Abstract: Recent video-based world models pair the scalability of autoregressive (AR) prediction with the visual quality of diffusion models. The choice of scene...
By Mikhail Dereviannykh, Vikram Voleti, Simon Donne, Mallikarjun Byrasandra Ramalinga Reddy, Shimon Vainer, Mark Boss
arXiv:2609.37925v1 Announce Type: cross
Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Trai...
By Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue
arXiv:2607. 21151v1 Announce Type: new Abstract: As Video Large Language Models are increasingly deployed in real-world applications, ensuring their safety alignment has become critical.
By Zhetong Zhang, Honghao Fu, Miao Xu, Yiwei Wang, Yujun Cai
arXiv:2609.39096v1 Announce Type: new
Abstract: Autoregressive video diffusion supports streaming generation and interactive control, but its KV cache grows with the generated history. Existing compr...
By Zeqi Xiao, Qingle Liu, Kaiwen Zhang, Yifan Zhou, Zihan Ding, Xingang Pan
arXiv:2604.11082v2 Announce Type: replace
Abstract: Visual glitches in video games degrade player experience and perceived quality, yet manual quality assurance cannot keep pace with the growing test...
By Yakun Yu, Ashley Wiens, Adri\'an Barahona-R\'ios, Benedict Wilkins, Saman Zadtootaghaj, Nabajeet Barman, Cor-Paul Bezemer
Token-Budget Distillation (TBD) is a parameter‑efficient fine‑tuning framework that adapts video vision‑language models to a fixed token budget. It freezes the pretrained backbone, updates only LoRA adapters, and incorporates FlashVID visual token compression. TBD uses a dual‑path teacher‑student design with full‑token supervision and compressed student optimization, enabling the student to recover full‑token semantics while remaining efficient under aggressive token reduction.
By Xiaoyang Guo, Guoping Luo, Jusheng Zhang, Keze Wang, Wenhao Wang