PARK: Accurate Block Retrieval for Sparse Attention in Video Diffusion Transformers
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.
arXiv:2607. 03012v1 Announce Type: cross Abstract: Video Diffusion Transformers (VDiTs) have demonstrated significant capabilities in high-fidelity video generation.
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
arXiv:2609.37001v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the compu...
arXiv:2609.15810v1 Announce Type: new Abstract: Diffusion Transformers deliver state-of-the-art video generation, but their long spatiotemporal sequences make attention the dominant deployment cost,...
ClusterAttention is a training‑free technique that speeds up bidirectional attention by recursively clustering keys and queries into fixed‑size, power‑of‑two blocks, enabling block‑sparse attention to match dense attention latency on GPUs. The method derives error bounds for sparse attention, showing tighter clusters can reduce error when compensated via centroids, and demonstrates significant speedups—up to six‑fold on large tabular data and 1.8× on video generation—while preserving over 99% of dense accuracy.