arXiv:2608. 09637v1 Announce Type: cross Abstract: Diffusion models have enabled high-quality video generation in recent years, but the high cost of iterative sampling hinders their practical deployment.
By Zian Li, Litong Gong, Borui Liao, Pengfei Liu, Xinyu Wang, Xinyuan Wei, Yifan Gao, Tiezheng Ge, Muhan Zhang
CrossDistill is a trajectory-level hybrid few-step distillation framework for diffusion models that balances quality and diversity by splitting the sampling trajectory at a crossover point. The high-noise interval uses a trajectory-preserving objective to maintain global mode coverage, while the low-noise interval applies a distribution-matching objective to sharpen local details, with the two stages coupled through the crossover state. This noise-level scheduling policy, demonstrated on text-to-video and image-to-video diffusion models, expands the few-step quality-diversity frontier by preserving seed-level variation while achieving competitive visual fidelity.
By Yuxi Liu, Haoyu Li, Yixiang Cai, Tengxu Sun, Zekun Zhang, Baole Ai, Ang Wang, Jiamang Wang, Lin Qu, Kun Yuan, Kai Zhang
The paper introduces DM-Align, a single-stage optimization framework that jointly performs distribution matching for distillation and aligns video generative models with human preferences. By deriving complementary gradient directions—one minimizing the gap between real and fake models and another guiding the model toward preferred samples—the method eliminates the need for separate reinforcement learning and distillation stages. Experiments on multiple foundational video models show that this sample-guided approach consistently outperforms both standalone variants and traditional two-stage pipelines.
By Jiuzhou Lin, Junlong Wu, Fei Zuo, Huan Ouyang, Dewen Fan, Boheng Zhang, Huaiqing Wang, Jia Sun, Fan Yang, Houde Liu, Kehai Chen, Min Zhang, Tingting Gao, Han Li
arXiv:2601. 09881v2 Announce Type: replace-cross Abstract: Large video diffusion and flow models have achieved remarkable success in high-quality video generation, but their use in real-time interactive applications remains limited due to their inefficient multi-step sampling process.
By Weili Nie, Julius Berner, Nanye Ma, Chao Liu, Saining Xie, Arash Vahdat
The paper introduces Uncertainty DMD, a lightweight framework that injects uncertainty into few-step autoregressive video distillation to counteract diversity collapse. By perturbing the first chunk’s timestep and employing a stochastic cache-writing mechanism for subsequent chunks, the method restores stochasticity without altering the model architecture. Experiments demonstrate consistent improvements in video diversity and motion dynamics while preserving visual quality.
By Zixuan Duan, Xunzhi Xiang, Yabo Chen, Xin Zhang, Changhan Liu, Haibin Huang, Chi Zhang, Qi Fan, Xuelong Li
The paper introduces Routed Forcing, a method that improves audio‑driven streaming avatar generation by selectively applying different distillation objectives to semantic regions and noise stages. It uses Data‑Forcing Distillation on person regions at high noise levels to restore motion diversity, while retaining Distribution Matching Distillation for mouth and background to keep lip sync and scene stability. Experiments show up to 45% better dynamics and 7–25% higher diversity compared to the previous Self Forcing approach.
By Zihan Su, Siwen Lu, Junhao Zhuang, Zeyue Xue, Haoyang Huang, Guanghao Li, Xiaofeng Tan, Chun Yuan, Nan Duan
ViRDM is a new post‑training method for few‑step causal video generation that eliminates the need for a large teacher model and an online critic. By applying representation distribution matching (RDM) with a precomputed target distribution, a lightweight VAE decoder, and staged vector–Jacobian products, ViRDM overcomes memory, optimization, and temporal dynamics challenges. The approach reduces GPU memory usage and training time, achieving state‑of‑the‑art VBench performance with only 20 generator updates and 16 A100 GPU‑hours.
By Zichong Meng, Chongjian Ge, Chun-Hao P. Huang, Yang Zhou, Huaizu Jiang
arXiv:2608.24674v1 Announce Type: new
Abstract: Joint text-to-video-audio generation produces synchronized visual and acoustic content, but the long sampling trajectories and heterogeneous multimodal...
By Xiaoda Yang, Yuxiang Liu, Kaiwen Zheng, Yuan Liu, Yibo Lai, Shengpeng Ji, Kai Jiang, Jianfei Chen, Xiaobin Hu, Shuicheng Yan, Jintao Zhang, Jun Zhu, Zhou Zhao
Few-step autoregressive (AR) video diffusion enables low-latency streaming generation, but existing post-training methods predominantly rely on Distribution Matching Distillation (DMD), requiring both...
The paper introduces DART, a training‑free technique that combines low‑rank coordinate transport with target‑schedule response calibration to improve the reuse of LoRA adapters in few‑step video diffusion models. By avoiding source training videos and using forward evaluations, DART raises the joint quality score on a four‑step Wan2.2 target from 0.9029 to 0.9227 and shifts macro functional retention from negative to positive. Component analysis shows that calibration drives most of the quality gains, while coordinate transport adds complementary benefits, and the method demonstrates consistent improvements across additional targets.
arXiv:2609.37925v1 Announce Type: cross
Abstract: Autoregressive (AR) video diffusion enables low-latency, streamable video generation, but prediction errors often accumulate over long rollouts. Trai...
By Chenjian Gao, Zhihao Hu, Jianqi Ma, Jun Zhang, Weidong Zhang, Tianfan Xue
The paper introduces HetA-DiT, a heterogeneous attention mechanism for video diffusion models that allocates computation based on token difficulty. A lightweight uncertainty branch predicts denoising difficulty, routing uncertain tokens through dense global attention while applying efficient local attention to reliable tokens. This adaptive routing retains global context where needed, offers a single parameter to balance quality and efficiency, and achieves competitive generation quality while only about 20% of tokens use dense attention.
By Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian