arXiv:2606. 16353v1 Announce Type: cross Abstract: Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets.
By Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai
arXiv:2608.27881v1 Announce Type: new
Abstract: Recently, many streaming video understanding methods have been proposed by constructing an external memory to store historical data for computational r...
By Yuxin Liu, Peiqin Zhuang, Yali Wang
arXiv:2609.00291v1 Announce Type: new
Abstract: Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily...
By Ce Zhang, Jing Bi, Jinxi He, Jianshu Zhang, Jingyang Lin, Yunzhong Xiao, Minghao Fu, Yaqi Xie, Zhentao Xie, Weicong Chen, Katia Sycara, Ming Zhou
The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding. It replaces the traditional store‑and‑retrieve approach with a retrieve‑and‑internalize strategy, organizing visual history into short, mid, and long‑term levels and progressively expanding memory receptive fields to internalize evidence into a compact latent memory. The method also employs confidence‑guided optimization to refine memory tokens, achieving state‑of‑the‑art results on online and offline video benchmarks.
The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.
By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
arXiv:2606. 17798v1 Announce Type: cross Abstract: Despite the remarkable progress of Video Large Language Models (Video-LLMs), current online architectures still struggle to simultaneously process continuous video streams, decide autonomously when to respond, and preserve long-horizon contextual memory.
By Zhenyu Yang, Kairui Zhang, Bing Wang, Shengsheng Qian, Changsheng Xu
The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.
By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv:2609.37559v1 Announce Type: new
Abstract: To serve as real-world personal assistants, streaming video models need persistent memory that retains past experiences for later use. Yet existing str...
By Jianguo Huang, Jinming Liu, Qiyao Wang, Liang Xu, Jianhang Li, Zhimian Wen, Mingda Li, Shule Lu, Zhicheng Wang, Yuhan Guo, Xin Jin, Wenjun Zeng
arXiv:2606. 26762v1 Announce Type: cross Abstract: Streaming video understanding (SVU) must answer queries that arrive asynchronously while visual tokens stream continuously under strict GPU-memory and query-time latency budgets.
By Le Tu Ngoc Minh (KAIST), Jinyeong Lim (KAIST), Dongsu Han (KAIST)
arXiv:2510. 09608v2 Announce Type: replace-cross Abstract: Vision-language models (VLMs) could power real-time assistants and autonomous agents, but they face a critical challenge: understanding near-infinite video streams without escalating latency and memory usage.
By Ruyi Xu, Guangxuan Xiao, Yukang Chen, Liuning He, Yao Lu, Song Han
arXiv:2609.23601v1 Announce Type: new
Abstract: Long-video understanding must capture transient visual evidence under strict token budgets, yet existing methods compress frames, append memory tokens,...
By Siru Zhong, Qiongyan Wang, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
CoFiE introduces a two‑stage evidence selection framework for streaming video understanding, separating a coarse, query‑agnostic filtering of visually distinctive frames from a fine, query‑specific refinement during LLM prefill. By filtering out redundant frames before expensive vision encoding, CoFiE reduces end‑to‑end latency while maintaining high accuracy. The method achieves state‑of‑the‑art performance on benchmarks such as StreamingBench and OvO‑Bench, improving accuracy by up to 3.15% and inference speed by up to 2.54× compared to prior approaches.
By Jing Jiang, Yiran Ling, Ruonan Li, Dimitrios Stamoulis, Jie Liu