arXiv AI By Wenxuan Miao, Haosong Liu, Weiming Hu, Zihan Liu, Aiyue Chen, Jianlin Yu, Yiwu Yao, Yiming Gan, Jieru Zhao, Jingwen Leng, Minyi Guo, Yu Feng

Kaleido: Algorithm-Hardware Co-Design for Video Diffusion Transformers by Exploiting Latent Space Correlations

Read the original on arXiv AI →

arXiv:2607. 13770v1 Announce Type: cross Abstract: Video diffusion transformers (vDiTs) generate high quality video but introduce extremely high compute cost due to the long diffusion timesteps and self attention computation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jun 21

Sol Video Inference Engine: Agent-Native Full-Stack Acceleration Framework for Efficient Video Generation

Modern video diffusion models achieve higher generation quality through scaling, but this also increases inference cost. Although many acceleration methods have been proposed, a central challenge is that the most effective acceleration strategy is highly instance-specific: a recipe that works well for one combination of model, hardware, and inference configuration often does not transfer to another.

arXiv Computer Vision
4d ago

FastVR: Efficient Streaming Video Restoration with One-Step Diffusion

FastVR is a streaming video restoration framework that uses a one‑step diffusion model to achieve strong restoration quality and temporal consistency while processing 1080p video at 11 FPS on a single H20 GPU. It addresses efficiency bottlenecks by combining a lightweight VAE with chunk‑wise causal attention, and improves inference speed and restoration quality through velocity consistency regularization and continuous trajectory learning during training. Experiments demonstrate that FastVR outperforms diffusion baselines in efficiency and achieves state‑of‑the‑art performance on both synthetic and real‑world benchmarks.

By Xiaoxu Chen, Qin Yang, Haoran Bai, Sibin Deng, Ying Chen