arXiv Machine Learning By Yuxi Liu, Yipeng Hu, Zekun Zhang, Kunze Jiang, Kun Yuan

Mixture of Distributions Matters: Dynamic Sparse Attention for Efficient Video Diffusion Transformers

Read the original on arXiv Machine Learning →

arXiv:2601. 11641v3 Announce Type: replace-cross Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.

arXiv Computer Vision
6d ago

Where Compute Matters: Heterogeneous Attention for Efficient Video Diffusion

The paper introduces HetA-DiT, a heterogeneous attention mechanism for video diffusion models that allocates computation based on token difficulty. A lightweight uncertainty branch predicts denoising difficulty, routing uncertain tokens through dense global attention while applying efficient local attention to reliable tokens. This adaptive routing retains global context where needed, offers a single parameter to balance quality and efficiency, and achieves competitive generation quality while only about 20% of tokens use dense attention.

By Olga Zatsarynna, Denis Korzhenkov, Juergen Gall, Amir Habibian, Mohsen Ghafoorian