Temporal-Attention Head Specialization During Video Diffusion Training
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2603.17825v2 Announce Type: replace Abstract: In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in vi...
arXiv:2605. 14513v2 Announce Type: replace-cross Abstract: Sparse attention accelerates video diffusion by allowing each attention head to focus on only a small subset of interactions.
arXiv:2609.08505v1 Announce Type: cross Abstract: Reliable video generation requires more than high-quality frames to form a coherent story: a model must maintain a persistent state, transporting vis...
arXiv:2601. 11641v3 Announce Type: replace-cross Abstract: While Diffusion Transformers (DiTs) have achieved notable progress in video generation, this long-sequence generation task remains constrained by the quadratic complexity inherent to self-attention mechanisms, creating significant barriers to practical deployment.
arXiv:2609.23658v1 Announce Type: cross Abstract: Despite impressive visual quality, state-of-the-art video diffusion models often generate content that violates real-world physical laws. While exist...
arXiv:2609.37001v1 Announce Type: cross Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the compu...