arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.
By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
By Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taix\'e, Despoina Paschalidou, Jonathan Lorraine, Aljo\v{s}a O\v{s}ep
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
By Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan
The paper introduces TetherMem, a training‑free, query‑aware memory router designed for streaming autoregressive video models that generate long videos in chunks. By separating subject and scene queries and modulating historical access with region‑ and age‑conditioned priors, TetherMem prevents the model from anchoring the scene to stale backgrounds and viewpoints, a problem termed memory‑anchored scene under‑progression. In blinded pairwise evaluations, TetherMem outperforms eight baseline methods in overall quality and scene progression, and on full 30‑second videos it maintains background, viewpoint, and scene changes while preserving subject identity and temporal continuity.
By Chen Li, Peng Zhang, Hanyu Zhou, Jialong Zuo, Fei Wang, Daiguo Zhou, Nong Sang, Changxin Gao
The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.
By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
RECAP-Forcing is a new method for long autoregressive video generation that addresses the memory challenge by organizing memory based on appearance novelty rather than recency. The approach retains key-value caches for newly appearing content—such as entering subjects, disoccluded regions, and new scenes—at the moment they first appear, ensuring consistent identities over time. It combines an attention sink for the initial scene with an optical-flow-based novelty bank for later frames, improving visual quality and semantic fidelity without adding learnable parameters.
By Haiyang Xu, Zheng Ding, Zhuowen Tu
arXiv:2607. 15271v1 Announce Type: cross Abstract: Online novel view synthesis from multi-view streaming videos faces a fundamental trade-off: maintaining a persistent, long-horizon memory to reconstruct temporarily occluded regions while operating under strict real-time constraints.
By Baback Elmieh, Lynn Tsai, Zeman Li, Srinivas Kaza, Tiancheng Sun, Gabor Csapo, Ali Behrouz, Yuan Deng, Stephen Lombardi, Steven M. Seitz, Xuan Luo
Ring Forcing is an autoregressive video diffusion framework that enhances long‑term memory by enforcing retrieval from distant history through a ring‑structured training strategy. It introduces a compression and timestep composition method to extend effective historical span to minutes, and a sparse RoPE mechanism for scalable memory adaptation. Experiments show that Ring Forcing outperforms state‑of‑the‑art models in minutes‑long coherence and object permanence.
By Bowen Xue, Brandon Y. Feng, Chenguo Lin, Yuchen Lin, Yujia Zeng, Lvmin Zhang, Maneesh Agrawala, Honglei Yan, Panwang Pan
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
By Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
arXiv:2608.27529v1 Announce Type: new
Abstract: Streaming 3D reconstruction from extremely long videos requires estimating camera motion and scene geometry online under bounded memory and computation...
By Jiarong Han, Jincheng Xiong, Yuzhou Liu, Linzhe Shi, Changjie Wu, Ning Guo, Mu Xu, Hang Zhang, Ming Qian
R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.
By Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao