arXiv:2608. 13043v1 Announce Type: new Abstract: Diffusion models have achieved dominant performance in visual generation but suffer from substantial inference overhead.
By Xichen Ye, Yifan Wu, Zhikang Xie, Xiangyu Yue, Cheng Jin, Weizhong Zhang
arXiv:2607. 29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising.
By Zhikang Xie, Xichen Ye, Yifan Wu, Haoshen Yu, Li chenan, Peizhu Gong, Weizhong Zhang, Cheng Jin
arXiv:2602.24208v2 Announce Type: replace-cross
Abstract: Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequenti...
By Yasaman Haghighi, Alexandre Alahi
arXiv:2608.29264v1 Announce Type: new
Abstract: Diffusion-based visual generative models deliver strong image and video synthesis quality but incur high inference costs because sequential samplers re...
By Yuhan Liu, Zongwei Hong, Jinglun Li, Linze Li, Shen Zhang, Yao Tang
arXiv:2609.37831v1 Announce Type: new
Abstract: Real-time diffusion-based video super-resolution (VSR) is in high demand for online streaming, yet stringent latency requirements often compromise gene...
By Xijun Wang, Xin Li, Suhang Yao, Zirui Lang, Bingchen Li, Zhibo Chen
arXiv:2607. 23159v1 Announce Type: new Abstract: Test-time search lets small video diffusion models rival larger ones, but costs 2-10x more.
By Shreshth Saini, Neil Birkbeck, Yilin Wang, Balu Adsumilli, Alan C. Bovik
arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.
By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh
arXiv:2609.39343v1 Announce Type: new
Abstract: Diffusion caching accelerates generation by replacing transformer computation with cached or predicted features at selected denoising steps. We introdu...
By Dong Wang, Wenwu Tang, Francesco Corti, Yun Cheng, Lothar Thiele, Olga Saukh
arXiv:2606. 26778v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs.
By Xuyue Huang, Zhe Chen, Wang Shen, Xiao-Ping Zhang
arXiv:2608. 05358v1 Announce Type: new Abstract: Federated learning repeatedly incurs local optimization and model-update transmission.
By Rahil Aftab, Vineet Kumar Rakesh, Soumya Mazumdar, Tapas Samanta
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.