Hugging Face Trending Papers

SCOPE: Subspace Clustering with Online Per-Head Top-K Estimation for Sparse Video Attention

Read the original on Hugging Face Trending Papers →

Diffusion Transformers (DiTs) incur quadratic self-attention cost over spatiotemporal tokens. Existing training-free sparse attention methods often construct sparse masks from block-level or cluster-level proxy scores, which can obscure fine-grained differences among keys and miss high contribution keys under aggressive sparsity.

Summary generated by The Flow from the publisher's feed. The full article lives at Hugging Face Trending Papers.