Ring Forcing is an autoregressive video diffusion framework that enhances long‑term memory by enforcing retrieval from distant history through a ring‑structured training strategy. It introduces a compression and timestep composition method to extend effective historical span to minutes, and a sparse RoPE mechanism for scalable memory adaptation. Experiments show that Ring Forcing outperforms state‑of‑the‑art models in minutes‑long coherence and object permanence.
By Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
By Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
arXiv:2609.37001v1 Announce Type: cross
Abstract: Diffusion Transformers (DiTs) enable high-quality video generation but suffer from substantial inference latency, primarily attributable to the compu...
By Xingyu Jia, Baole Ai, Ang Wang, Kang Zhao, Yong Li
High-resolution image and video diffusion models, including SD3, FLUX, and recent video diffusion transformers, have substantially improved generative quality but remain expensive at inference time because they repeatedly evaluate attention-heavy denoisers over many sampling steps. We address this inefficiency by exploiting redundancy in intermediate diffusion features rather than changing model weights or retraining.
arXiv:2509.25998v4 Announce Type: replace
Abstract: In light of recent progress in video editing, deep learning models focusing on both spatial and temporal dependencies have emerged as the primary m...
By Abdelilah Aitrouga, Youssef Hmamouche, Amal El Fallah Seghrouchni
Video DeltaNet (VDN) introduces a hybrid attention mechanism for livestream video generation, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). VDA updates memory once per frame, integrating spatial tokens, while separate output projections and learnable gates balance the two branches. Applied to MiniMax H3, VDN achieves a 14.5× speedup over the dense baseline, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs.
By Haocheng Xi, Yiming Xie, Hexu Zhao, Yiwen Zhang, Michael Liu, Thomas Creavin, Kurt Keutzer, Xiuyu Li, Zhaoyang Lv, Chenfeng Xu, Haiwen Feng
arXiv:2609.34895v2 Announce Type: replace
Abstract: Existing online video segmentation methods struggle to track objects in long, complex videos with long-term occlusions. We hypothesize that this li...
By Narges Norouzi, Niccol\`{o} Cavagnero, Idil Esen Zulfikar, Bastian Leibe, Gijs Dubbelman, Daan de Geus
arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.
By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.
By Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
Video DeltaNet (VDN) introduces a hybrid attention mechanism for video diffusion models, combining local Softmax attention with a bidirectional linear memory branch called Video Delta Attention (VDA). The design updates memory once per frame, uses separate output projections and learnable gates to balance the two branches, and employs a staged teacher‑alignment recipe to integrate the new pathway into pretrained models. When applied to MiniMax H3, VDN achieves a 14.5× speedup, completing 14.3‑second, 768p video denoising in 6.70 seconds on eight NVIDIA B200 GPUs compared to the 50‑step dense baseline.
arXiv:2512.07480v2 Announce Type: replace
Abstract: While traditional and neural video codecs (NVCs) have achieved remarkable rate-distortion performance, improving perceptual quality at low bitrates...
By Naifu Xue, Zhaoyang Jia, Jiahao Li, Bin Li, Zihan Zheng, Yuan Zhang, Yan Lu
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.