SVMemAgent introduces a streaming video memory (SVMem) that continuously updates a compact representation of observed frames for online keyframe selection without prior knowledge of video length, query, or future frames. The agent decides at each timestep whether to replace an existing memory frame with a new one or discard it, trained via Group Relative Policy Optimization using task-driven rewards from question-answer pairs. Experiments demonstrate that SVMemAgent outperforms existing online baselines and rivals offline methods, and its learned policy tends to favor frames containing textual information, potentially aiding downstream VideoQA tasks.
By Dohwan Ko, Ji Soo Lee, Pierce Chuang, Debojeet Chatterjee, Ashish Shenoy, Yichao Lu, Seungwhan Moon, Xin Luna Dong, Vikas Bhardwaj, Hyunwoo J. Kim
arXiv:2606. 16353v1 Announce Type: cross Abstract: Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets.
By Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai
arXiv:2609.00291v1 Announce Type: new
Abstract: Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily...
By Ce Zhang, Jing Bi, Jinxi He, Jianshu Zhang, Jingyang Lin, Yunzhong Xiao, Minghao Fu, Yaqi Xie, Zhentao Xie, Weicong Chen, Katia Sycara, Ming Zhou
arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.
By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin
The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding. It replaces the traditional store‑and‑retrieve approach with a retrieve‑and‑internalize strategy, organizing visual history into short, mid, and long‑term levels and progressively expanding memory receptive fields to internalize evidence into a compact latent memory. The method also employs confidence‑guided optimization to refine memory tokens, achieving state‑of‑the‑art results on online and offline video benchmarks.
The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.
By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.
By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
The paper introduces Caption‑once, Frames‑on‑Demand (CFD), a budget‑aware edge‑cloud framework for long‑video understanding. CFD first runs a single offline captioning pass on the edge to build a dual‑track narrative index—an event‑level story skeleton and a clip‑level micro‑log—that is cached for future queries. At query time, a cloud‑side MLLM uses a Visual‑Need Router to decide whether to retrieve keyframes for perceptual questions, thereby limiting visual processing while preserving temporal structure in language space.
By Weitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie, Shan Gao, Jiankang Deng, Songcen Xu, Jifei Song, Zhensong Zhang
arXiv:2506. 01274v2 Announce Type: replace-cross Abstract: Recent progress in Large Multi-modal Models (LMMs) has enabled effective vision-language reasoning, yet the ability to video understanding remains constrained by suboptimal frame selection strategies, albeit with the rapid development of video-specialized LMMs.
By Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro
arXiv:2602. 01801v2 Announce Type: replace-cross Abstract: Autoregressive video diffusion models enable streaming generation, opening the door to long-form synthesis, video world models, and interactive neural game engines.
By Dvir Samuel, Issar Tzachor, Matan Levy, Michael Green, Gal Chechik, Rami Ben-Ari
arXiv:2608. 08612v1 Announce Type: cross Abstract: Recently, retrieval-augmented and memory-augmented methods have emerged as two promising paradigms for long-video question answering.
By Caijun Yan, Yang Zhou, Meixing Shi, Haoran Sun, Yichen Li, Yuxiang Cai, Yankai Jiang
EventMemAgent is an active online video agent that uses a hierarchical memory module to handle continuous perception and long‑range reasoning in streaming video. The framework employs a short‑term memory layer to detect event boundaries and sample frames within a fixed buffer, while a long‑term memory layer archives observations event‑by‑event. It also incorporates a multi‑granular perception toolkit and Agentic Reinforcement Learning to internalize reasoning and tool‑use strategies, achieving competitive results on online video benchmarks.
By Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu