arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.
By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
By Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taix\'e, Despoina Paschalidou, Jonathan Lorraine, Aljo\v{s}a O\v{s}ep
Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content. However, extending these models to minute-level generation remains challenging: the limited KV-cache budget prevents the model from retaining the full history, while repeatedly conditioning on self-generated frames induces a context distribution shift that accumulates over time, leading to visual artifacts, quality degradation, and temporal drift.
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.
By Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan
The paper introduces TetherMem, a training‑free, query‑aware memory router designed for streaming autoregressive video models that generate long videos in chunks. By separating subject and scene queries and modulating historical access with region‑ and age‑conditioned priors, TetherMem prevents the model from anchoring the scene to stale backgrounds and viewpoints, a problem termed memory‑anchored scene under‑progression. In blinded pairwise evaluations, TetherMem outperforms eight baseline methods in overall quality and scene progression, and on full 30‑second videos it maintains background, viewpoint, and scene changes while preserving subject identity and temporal continuity.
By Chen Li, Peng Zhang, Hanyu Zhou, Jialong Zuo, Fei Wang, Daiguo Zhou, Nong Sang, Changxin Gao
The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.
By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras