Why Do Video Diffusion Models Violate Physics? Unveiling the Flaws in Attention Mechanisms
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2609.36832v1 Announce Type: new Abstract: Text-to-video (T2V) diffusion models can generate realistic depictions of actions such as kicking, stabbing, and shooting, raising safety concerns that...
arXiv:2606. 05328v1 Announce Type: cross Abstract: Modern video diffusion models generate increasingly realistic and temporally coherent videos, motivating their use as candidate world simulators.
arXiv:2603.17825v2 Announce Type: replace Abstract: In this work, we study the role of Massive Activations (MAs), which are rare, high-magnitude spikes confined to a few fixed hidden dimensions in vi...
arXiv:2603. 16870v3 Announce Type: replace-cross Abstract: Recent advances in video generation have revealed an unexpected phenomenon: diffusion-based video models exhibit non-trivial reasoning capabilities.
arXiv:2609.31654v2 Announce Type: replace-cross Abstract: Video diffusion transformers depend on temporal attention to coordinate information across frames, yet nearly everything known about this mec...
arXiv:2605. 19398v3 Announce Type: replace-cross Abstract: Image-to-video models often generate videos that remain overly static, compared to text-to-video models.