arXiv AI By Yuta Oshima, Shohei Taniguchi, Masahiro Suzuki, Yutaka Matsuo

SSM Meets Video Diffusion Models: Efficient Long-Term Video Generation with Structured State Spaces

Read the original on arXiv AI →

arXiv:2403. 07711v5 Announce Type: replace-cross Abstract: Given the remarkable achievements in image generation through diffusion models, the research community has shown increasing interest in extending these models to video generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computer Vision
Aug 28

Ring Forcing: Towards Precise Long-Term Memory for Autoregressive Video Diffusion

Ring Forcing is an autoregressive video diffusion framework that enhances long‑term memory by enforcing retrieval from distant history through a ring‑structured training strategy. It introduces a compression and timestep composition method to extend effective historical span to minutes, and a sparse RoPE mechanism for scalable memory adaptation. Experiments show that Ring Forcing outperforms state‑of‑the‑art models in minutes‑long coherence and object permanence.

By Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan
Hugging Face Trending Papers
Jul 26

OmniCache: Multidimensional Hierarchical Feature Caching For Diffusion Models

High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive at inference time because they repeatedly evaluate attention-heavy denoisers over many sampling steps. We address this inefficiency by exploiting redundancy in intermediate diffusion features rather than changing model weights or retraining.

arXiv Machine Learning
Sep 18

Video DeltaNet: A Video-Native Hybrid Attention for Livestream Video Generation

Video DeltaNet (VDN) introduces a hybrid attention mechanism for livestream video generation, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). VDA updates memory once per frame, integrating spatial tokens, while separate output projections and learnable gates balance the two branches. Applied to MiniMax H3, VDN achieves a 14.5× speedup over the dense baseline, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs.

By Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng