arXiv AI

S-EMBER: A Large-Scale Benchmark for Streaming Egocentric Memory Retrieval

arXiv:2607. 02689v1 Announce Type: cross Abstract: As wearable devices enable continuous first-person recording, AI assistants must reason across long time horizons to recall past experiences-a capability known as episodic memory.

arXiv AI
Aug 20

EgoMemReason: A Memory-Driven Reasoning Benchmark for Long-Horizon Egocentric Video Understanding

EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.

By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv AI
Sep 17

CapMem: A Benchmark for Caption-Based Episodic Memory in Egocentric Video

CapMem is a new benchmark for evaluating caption-based episodic memory in egocentric video. It contains 75 videos (33.7 hours total) and 1,000 multiple-choice questions across 16 scenarios, designed to test the Episodic Memory Video Caption QA task. Experiments show that using captions as memory outperforms direct VideoQA on long videos, and a caption-guided retrieve-and-verify approach further boosts accuracy.

By Dingli Liang, Yiqiao Xie, Yukai Huang, Zhaokai Wang, Weitong Cai, Guangwen Feng, Jifei Song, Zhensong Zhang, Hang Zhang
arXiv AI
Aug 20

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

Event-Causal RAG (EC‑RAG) is a lightweight retrieval‑augmented framework designed for reasoning over ultra‑long and streaming videos. It segments video streams into semantically complete events using a dual visual‑audio sentinel mechanism, representing each event as a State‑Event‑State (SES) structure that captures pre‑event, event, and post‑event states. During question answering, bidirectional graph retrieval accesses relevant predecessor and successor events from a dual vector‑graph memory, and answers are generated using both this structured memory and the corresponding video evidence. The authors also introduce ECV‑1H, an hour‑scale long‑video QA benchmark with over 150 hours of untrimmed video and 1,251 human‑annotated QA pairs, where EC‑RAG achieves significant accuracy gains across multiple video foundation models while maintaining efficient streaming memory usage on a single RTX 5090 GPU.

By Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin
arXiv AI
Jun 8

Watch, Remember, Reason: Human-View Video Understanding with MLLMs

arXiv:2606. 07433v1 Announce Type: cross Abstract: Video understanding is being rapidly transformed by multimodal large language models (MLLMs), as research moves from short clips to long, multimodal, and knowledge-intensive video scenarios.

By Jiahao Meng, Yue Tan, Qi Xu, Kuan Gao, Weisong Liu, Yanwei Li, Jason Li, Lingdong Kong, Haochen Wang, Qianyu Zhou, Jiangning Zhang, Guangliang Cheng, Yunhai Tong, Lu Qi, Minghsuan Yang
arXiv Machine Learning
Jul 7

Agentic Very Long Video Understanding

arXiv:2601. 18157v3 Announce Type: replace-cross Abstract: The advent of always-on personal AI assistants, enabled by all-day wearable devices such as smart glasses, demands a new level of contextual understanding, one that goes beyond short, isolated events to encompass the continuous, longitudinal stream of egocentric video.

By Aniket Rege, Arka Sadhu, Yuliang Li, Kejie Li, Ramya Korlakai Vinayak, Yuning Chai, Yong Jae Lee, Hyo Jin Kim
arXiv AI
Sep 2

Learning to Remember: End-to-End Training of Memory Agents for Long-Context Reasoning

The paper introduces the Unified Memory Agent (UMA), a system that builds a query‑agnostic external memory from a data stream and reuses it across multiple question‑answering sessions. UMA employs a single policy to manage a structured Memory Bank via CRUD operations and uses Task‑Stratified GRPO to supervise memory maintenance based on QA trajectory rewards. The authors also present Ledger‑QA, a benchmark for long‑horizon state tracking, and demonstrate that UMA outperforms other methods on test‑time learning and accurate‑retrieval tasks, with UMA‑Specialist further improving performance after task adaptation.

By Kehao Zhang, Shangtong Gui, Sheng Yang, Wei Chen, Yang Feng