arXiv AI

MemLife: Curating and Reasoning over Long-Term Egocentric Video Memories

arXiv AI
Aug 20

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.

By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv AI
Jul 15

ReflectWorld-MM: An Entity-Oriented Multimodal Memory System for Open-Ended Video Streams

arXiv:2607. 09759v2 Announce Type: replace-cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
arXiv Computer Vision
Sep 14

EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use

EventMemAgent is an active online video agent that uses a hierarchical memory module to handle continuous perception and long‑range reasoning in streaming video. The framework employs a short‑term memory layer to detect event boundaries and sample frames within a fixed buffer, while a long‑term memory layer archives observations event‑by‑event. It also incorporates a multi‑granular perception toolkit and Agentic Reinforcement Learning to internalize reasoning and tool‑use strategies, achieving competitive results on online video benchmarks.

By Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu
arXiv Computer Vision
Sep 4

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.

By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
arXiv Computer Vision
Sep 7

ICM-Bench: Person-Level Identity Reasoning in Multimodal Agents with Long-Term Memory

ICM-Bench is a new benchmark for evaluating identity-centric reasoning in multimodal agents with long-term memory. It consists of 839 synthetic video clips totaling 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark isolates the ability to maintain recurring person identities and reason over their cross-time relations, and compares various baseline systems, showing that while Gemini 3.1 Pro performs well overall, its accuracy drops on questions requiring long-term identity profiles.

By Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen
arXiv AI
Jul 14

ReflectWorld-MM: An Entity-Oriented Multi-Media Memory System for Open-Ended Video Streams

arXiv:2607. 09759v1 Announce Type: cross Abstract: Building assistants that can continually watch the world, remember what they see, and reason over their accumulated experience is a long-standing goal, and recently multimodal agents equipped with long-term memory over video streams have attracted increasing interest.

By Xiaokang Ma, Yifan Sun, Zhihong Jin, Jie Gu, Yudong Luo, Shenyi Shao, Chu Tang, Jingmin Chen, Li Pu
arXiv AI
Sep 10

MemForest: Efficient Agent Memory Management via EventTree Partitioning and Progressive Merging

arXiv:2609.08273v1 Announce Type: new Abstract: Agent memory systems have demonstrated significant potential in long-term dialogue, personalized assistants, and video understanding. However, continuo...

By Junxi Wang, Te Sun, Jiayi Zhu, Chen Zhang, Siyuan Li, Xuyang Liu, Zichen Wen, Xiaobing Tu, Jinkui Ren, Xiantao Zhang, Ziqi Yuan, Linfeng Zhang
Hugging Face Trending Papers
Sep 3

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding. It replaces the traditional store‑and‑retrieve approach with a retrieve‑and‑internalize strategy, organizing visual history into short, mid, and long‑term levels and progressively expanding memory receptive fields to internalize evidence into a compact latent memory. The method also employs confidence‑guided optimization to refine memory tokens, achieving state‑of‑the‑art results on online and offline video benchmarks.