arXiv Computer Vision

The Past Frames the Future: Memory for Autoregressive Video Generation

arXiv Computer Vision
Aug 31

Relax Forcing: Relaxed KV-Memory for Consistent Long Video Generation

The paper introduces Relax Forcing, a training‑free memory mechanism for autoregressive video diffusion that structures temporal context into Sink, Tail, and History frames. By selecting History frames with a relaxation criterion, the method reduces error accumulation and attention overhead while preserving motion dynamics. Experiments on VBench‑Long demonstrate that this structured memory improves long‑video generation quality over existing baselines.

By Zengqun Zhao, Yanzuo Lu, Ziquan Liu, Jifei Song, Jiankang Deng, Ioannis Patras
arXiv Machine Learning
Jun 9

Echo-Memory: A Controlled Study of Memory in Action World Models

arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.

By Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan
arXiv Computer Vision
Aug 31

LayerRecall: A State-Conditioned Memory Router for Long-Horizon Consistency in Video Generation

LayerRecall is a memory router for autoregressive video diffusion that selectively retrieves and injects historical key/value states into specific layers of the model, based on the current context. It addresses the problem that existing memory mechanisms expose nonlocal history but do not guarantee effective use, by recognizing that different layers prefer current, recent, or distant context. The method, combined with Cross‑Horizon Prediction Matching, achieves state‑of‑the‑art long‑range consistency on MemoBench and MovieBench while maintaining local continuity and incurring negligible inference overhead.

By Yixuan Ding, Jiahao Kong, Wei Huang, Ruijie Quan, Yi Yang
arXiv AI
1d ago

Compress to Remember: Learning Compact Memory via On-Policy Distillation for Long Video Generation

The paper introduces Prediction‑Aligned Context Compaction (PACC), a method that learns a compact memory representation for long‑video generation by distilling a frozen video generator. PACC trains a compressor to aggregate past frames into memory tokens, using the generator as both teacher and student during on‑policy distillation. Experiments on MBench and VBench‑Long show that PACC improves memory‑event coverage and consistency, achieving better scores than strong baselines and producing competitive minute‑long videos.

By Xiaoyu Wu, Weihang Guo, Yifei Wang, Xinze Feng, Lydia E. Kavraki, Zhiwei Steven Wu
arXiv Computer Vision
1d ago

APM-Bench: Benchmarking Cross-session Persistent Memory for Egocentric Streaming Video Assistants

arXiv:2609.37559v1 Announce Type: new Abstract: To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing str...

By Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, Zhicheng Wang, Yuhan Guo, Xin Jin, Wenjun Zeng
arXiv Computer Vision
Sep 4

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.

By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
arXiv AI
Aug 20

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.

By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv AI
Jun 12

TetherCache: Stabilizing Autoregressive Long-Form Video Generation with Gated Recall and Trusted Alignment

arXiv:2606. 13035v1 Announce Type: cross Abstract: Autoregressive video diffusion models provide a natural formulation for streaming and variable-length video generation by conditioning newly generated frames on previously generated content.

By Yu Meng, Xiangyang Luo, Letian Li, Wenyuan Jiang, Chen Gao, Xinlei Chen, Yong Li, Xiao-Ping Zhang
arXiv AI
Jul 14

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

arXiv:2607. 09759v1 Announce Type: cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu