arXiv:2607. 29398v1 Announce Type: new Abstract: Diffusion models have revolutionized generative tasks but incur high latency due to iterative denoising.
By Zhikang Xie, Xichen Ye, Yifan Wu, Haoshen Yu, Li chenan, Peizhu Gong, Weizhong Zhang, Cheng Jin
arXiv:2606. 26778v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) have driven substantial progress in image and video generation but suffer from prohibitive computational costs.
By Xuyue Huang, Zhe Chen, Wang Shen, Xiao-Ping Zhang
arXiv:2603.01623v2 Announce Type: replace
Abstract: Diffusion models have become the dominant tool for high-fidelity image and video generation, yet are critically bottlenecked by their inference spe...
By Jiaqi Han, Juntong Shi, Puheng Li, Haotian Ye, Qiushan Guo, Stefano Ermon
Diffusion models have achieved remarkable success in image and video generation, yet the high computational cost of iterative sampling remains a critical bottleneck for practical deployment. Feature c...
arXiv:2602.24208v2 Announce Type: replace-cross
Abstract: Diffusion models achieve state-of-the-art video generation quality, but their inference remains expensive due to the large number of sequenti...
By Yasaman Haghighi, Alexandre Alahi
The paper introduces TRACK, a training‑free trajectory routing method that accelerates video diffusion by selectively switching between large and small models during denoising steps. A calibration process generates a disagreement score map, guiding the selection of the appropriate model at each step to maintain quality while reducing computational cost. Experiments on Wan 2.1, Cosmos 3, TurboDiffusion, and FastVideo show speedups ranging from 1.95× to 2.73× with comparable quality and diversity.
By Mustafa Munir, Huy Vu, Shreyas Misra, Rohit Jena, Sajad Norouzi, Ali Taghibakhshi, Anis Ahmad, Anjul Patney, Pavlo Molchanov, Nima Tajbakhsh
Accelerating Video Diffusion via Training-Free Trajectory Routing (TRACK) introduces a heterogeneous denoising strategy that switches between large and small diffusion models at selected steps, determined by a calibration process that measures disagreement between model predictions. By routing quality-sensitive steps to the large model and low-disagreement steps to the small model, TRACK achieves significant speedups—up to 2.73×—across several video diffusion benchmarks while maintaining comparable quality and diversity. The method requires no retraining, architectural changes, or online dual-model evaluation, making it a practical acceleration paradigm for video diffusion.
arXiv:2412. 18911v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiT) have become the dominant methods in image and video generation yet still suffer substantial computational costs.
By Chang Zou, Shikang Zheng, Evelyn Zhang, Runlin Guo, Haohang Xu, Zhengyi Shi, Conghui He, Xuming Hu, Linfeng Zhang
arXiv:2604. 22901v2 Announce Type: replace Abstract: Diffusion models achieve remarkable success in time series generation.
By Dong Liu, Yanxuan Yu, Ying Nian Wu
arXiv:2606. 15615v1 Announce Type: new Abstract: Diffusion Transformers with Mixture-of-Experts (DiT-MoE) improve model capacity under sparse activation, but diffusion inference is still bottlenecked by redundant computation across timesteps.
By Maoliang Li, Haojing Chen, Jiayu Chen, Zihao Zheng, Xinhao Sun, Hailong Zou, Xiang Chen
High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive at inference time because they repeatedly evaluate attention-heavy denoisers over many sampling steps. We address this inefficiency by exploiting redundancy in intermediate diffusion features rather than changing model weights or retraining.
arXiv:2602. 13357v3 Announce Type: replace-cross Abstract: Diffusion Transformers (DiTs) achieve state-of-the-art performance in high-fidelity image and video generation but suffer from expensive inference due to their iterative denoising structure.
By Dong Liu, Yanxuan Yu, Ben Lengerich, Ying Nian Wu