arXiv AI

A neural network that maintains and retrieves memories based on context

A recurrent neural network with an episodic memory buffer was trained to infer situational context while watching naturalistic movies, using Bayesian inference to predict upcoming scenes. The inferred context modulates the network’s recurrent connectivity in a low‑rank manner, producing activity patterns that best match human fMRI responses. Context also guides episodic memory retrieval via a key‑value self‑attention system, enabling the model to retrieve memories faster and more accurately than a context‑agnostic version.

arXiv AI
Jul 28

Temporal Context Reinstatement Drives Episodic-Like Order Memory in Long-Context Language Models

arXiv:2607. 22575v1 Announce Type: new Abstract: Human episodic memory supports the retrieval of experiences that unfold over extended timescales, yet the computational mechanisms underlying this ability remain debated due to the limited mechanistic accessibility in long-term memory experiments in humans.

By Mathis Pink, Vy Ai Vo, Qinyuan Wu, Jianing Mu, Javier Turek, Uri Hasson, Kenneth A. Norman, Sebastian Michelmann, Alexander Huth, Mariya Toneva
arXiv Computer Vision
Sep 4

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.

By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
arXiv Computer Vision
Aug 28

RECAP-Forcing: Retaining Content Appearances for Long Video Generation

RECAP-Forcing is a new method for long autoregressive video generation that addresses the memory challenge by organizing memory based on appearance novelty rather than recency. The approach retains key-value caches for newly appearing content—such as entering subjects, disoccluded regions, and new scenes—at the moment they first appear, ensuring consistent identities over time. It combines an attention sink for the initial scene with an optical-flow-based novelty bank for later frames, improving visual quality and semantic fidelity without adding learnable parameters.

By Haiyang Xu, Zheng Ding, Zhuowen Tu
Hugging Face Trending Papers
Sep 3

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding. It replaces the traditional store‑and‑retrieve approach with a retrieve‑and‑internalize strategy, organizing visual history into short, mid, and long‑term levels and progressively expanding memory receptive fields to internalize evidence into a compact latent memory. The method also employs confidence‑guided optimization to refine memory tokens, achieving state‑of‑the‑art results on online and offline video benchmarks.

arXiv Machine Learning
Jun 9

Echo-Memory: A Controlled Study of Memory in Action World Models

arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.

By Wayne King, Zeyue Xue, Yuxuan Bian, Jie Huang, Haoran Li, Yaowei Li, Yaofeng Su, Yuming Li, Haoyu Wang, Shiyi Zhang, Songchun Zhang, Yuwei Niu, Sihan Xu, Junhao Zhuang, Haoyang Huang, Nan Duan
arXiv AI
Aug 20

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.

By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv AI
4d ago

ReMem: Rethinking Perception and Memory in Long-Context Recommendation Agents

ReMem is a new recommendation agent framework that rethinks perception and memory for long-context recommendation tasks. It replaces raw HTML parsing with OCR-based multimodal perception from screenshots, extracting structured information in a platform-agnostic way. The framework also introduces a chunk-wise sequential memory update strategy and a multi-memory GRPO variant to efficiently model evolving user preferences over arbitrarily long interaction histories, achieving a 5.16% average improvement over state-of-the-art baselines on three recommendation agent tasks.

By Haohao Qu, Yongcheng Jing, Chun Hin Chan, Shanru Lin, Wenqi Fan, Dacheng Tao
arXiv AI
Jun 8

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.

By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang