ICM-Bench is a new benchmark for evaluating identity-centric reasoning in multimodal agents with long-term memory. It consists of 839 synthetic video clips totaling 141 minutes and 1,217 open-ended questions about six recurring adults in a one-year life album. The benchmark isolates the ability to maintain recurring person identities and reason over their cross-time relations, and compares various baseline systems, showing that while Gemini 3.1 Pro performs well overall, its accuracy drops on questions requiring long-term identity profiles.
By Shidu Ren, Yunze Liu, Xing Liu, Chi-Hao Wu, Enmin Zhou, Junxiao Shen
EgoMemReason is a new benchmark for week‑long egocentric video understanding that focuses on memory‑driven reasoning rather than simple perception tasks. It tests three memory types—entity, event, and behavior—across 500 questions, each requiring evidence from an average of 5.1 video segments and 25.9 hours of backtracking. Evaluation of 17 models shows that even the best achieves only 39.6% accuracy, highlighting the difficulty of long‑horizon memory in multimodal systems.
By Ziyang Wang, Yue Zhang, Shoubin Yu, Ce Zhang, Zengqi Zhao, Jaehong Yoon, Hyunji Lee, Gedas Bertasius, Mohit Bansal
arXiv:2608. 08612v1 Announce Type: cross Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering.
By Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li, Yuxiang Cai, Yankai Jiang
EventMemAgent is an active online video agent that uses a hierarchical memory module to handle continuous perception and long‑range reasoning in streaming video. The framework employs a short‑term memory layer to detect event boundaries and sample frames within a fixed buffer, while a long‑term memory layer archives observations event‑by‑event. It also incorporates a multi‑granular perception toolkit and Agentic Reinforcement Learning to internalize reasoning and tool‑use strategies, achieving competitive results on online video benchmarks.
By Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu
Self-Evolving Multimedia Verification through Memory Consolidation of Contestation Experiences (SEMV) is a multi-agent framework that uses provenance-bearing arguments to link evidence, reasoning, human contestation, and memory. It integrates arena-based quantitative bipolar argumentation, causal and scoped revision, and verification-gated memory consolidation with explicit conflict retention. On the COSMOS benchmark, SEMV achieves 91.88% accuracy, reducing negative transfer from 5.7% to 0.2%, and on the CTR benchmark it corrects 96.7% of initial errors while saving 52.8% of compute.
By Truong Thanh Hung Nguyen, Vo Thanh Khang Nguyen, Hoang-Loc Cao, Phuc Ho, Truong Thinh Nguyen, Van Pham, Hung Cao
arXiv:2606. 05008v1 Announce Type: cross Abstract: As multi-modal models advance towards long-form video understanding, memory emerges as a critical capability.
By Jie Huang, Ruixun Liu, Sirui Sun, Xinyi Yang, Yin Li, Yixin Zhu, Yiwu Zhong