Hugging Face Trending Papers

SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection

arXiv Computer Vision
Sep 17

SVMemAgent: A Streaming Video Memory Agent for Query-Agnostic Online Frame Selection

SVMemAgent introduces a streaming video memory (SVMem) that continuously updates a compact representation of observed frames for online keyframe selection without prior knowledge of video length, query, or future frames. The agent decides at each timestep whether to replace an existing memory frame with a new one or discard it, trained via Group Relative Policy Optimization using task-driven rewards from question-answer pairs. Experiments demonstrate that SVMemAgent outperforms existing online baselines and rivals offline methods, and its learned policy tends to favor frames containing textual information, potentially aiding downstream VideoQA tasks.

By Dohwan Ko, Ji Soo Lee, Pierce Chuang, Debojeet Chatterjee, Ashish Shenoy, Yichao Lu, Seungwhan Moon, Xin Luna Dong, Vikas Bhardwaj, Hyunwoo J. Kim
arXiv Computer Vision
Sep 2

StreamScout: Learning When to Look Deeper for Streaming Video Understanding

arXiv:2609.00291v1 Announce Type: new Abstract: Streaming video understanding requires answering questions that arrive at arbitrary moments over an unbounded video stream. Existing systems primarily...

By Ce Zhang, Jing Bi, Jinxi He, Jianshu Zhang, Jingyang Lin, Yunzhong Xiao, Minghao Fu, Yaqi Xie, Zhentao Xie, Weicong Chen, Katia Sycara, Ming Zhou
arXiv AI
Jul 29

Reasoning with Memory: A Temporal Granularity-Adaptive Framework for Training-Free Long Video Understanding

arXiv:2607. 24794v1 Announce Type: new Abstract: While Multimodal Large Language Models (MLLMs) demonstrate superior generalization in fundamental video tasks, restricted context windows limit their long video understanding.

By Linghao Meng, Qiankun Li, Junyuan Mao, Pujin Liao, Zhicheng He, Enbo Zhang, Kun Wang, Yang Liu, Huazhu Fu, Yueming Jin
Hugging Face Trending Papers
Sep 3

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding. It replaces the traditional store‑and‑retrieve approach with a retrieve‑and‑internalize strategy, organizing visual history into short, mid, and long‑term levels and progressively expanding memory receptive fields to internalize evidence into a compact latent memory. The method also employs confidence‑guided optimization to refine memory tokens, achieving state‑of‑the‑art results on online and offline video benchmarks.

arXiv AI
1d ago

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.

By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv Computer Vision
Sep 4

Beyond Retrieval: Progressive Latent Memory Evolution for Streaming Video Understanding

The paper introduces LatentStream, a progressive latent working memory framework for streaming video understanding that replaces the traditional store‑and‑retrieve paradigm with a retrieve‑and‑internalize approach. It organizes visual history into short, mid, and long‑term levels using Jenks‑guided adaptive consolidation, then expands memory receptive fields to iteratively retrieve and internalize evidence into a compact latent memory. A confidence‑guided optimization further refines this memory, leading to state‑of‑the‑art performance on online and offline video benchmarks.

By Hongyu Qu, Guangming Yao, Ling Xing, Xiaobin Hu, Rongxing Ding, Guibin Zhang, Fan Zhang, Yi Yuan, Xiangbo Shu, Shuicheng Yan
arXiv Computer Vision
Sep 11

Caption-once, Frames-on-Demand: Visual-Need Routing for Budget-Aware Agentic Long Video Understanding

The paper introduces Caption‑once, Frames‑on‑Demand (CFD), a budget‑aware edge‑cloud framework for long‑video understanding. CFD first runs a single offline captioning pass on the edge to build a dual‑track narrative index—an event‑level story skeleton and a clip‑level micro‑log—that is cached for future queries. At query time, a cloud‑side MLLM uses a Visual‑Need Router to decide whether to retrieve keyframes for perceptual questions, thereby limiting visual processing while preserving temporal structure in language space.

By Weitong Cai, Hang Zhang, Yukai Huang, Yiqiao Xie, Shan Gao, Jiankang Deng, Songcen Xu, Jifei Song, Zhensong Zhang
arXiv AI
Jun 12

ReFoCUS: Reinforcement-guided Frame Optimization for Contextual Understanding

arXiv:2506. 01274v2 Announce Type: replace-cross Abstract: Recent progress in Large Multi-modal Models (LMMs) has enabled effective vision-language reasoning, yet the ability to video understanding remains constrained by suboptimal frame selection strategies, albeit with the rapid development of video-specialized LMMs.

By Hosu Lee, Junho Kim, Hyunjun Kim, Yong Man Ro
arXiv Computer Vision
Sep 14

EventMemAgent: Hierarchical Event-Centric Memory for Online Video Understanding with Adaptive Tool Use

EventMemAgent is an active online video agent that uses a hierarchical memory module to handle continuous perception and long‑range reasoning in streaming video. The framework employs a short‑term memory layer to detect event boundaries and sample frames within a fixed buffer, while a long‑term memory layer archives observations event‑by‑event. It also incorporates a multi‑granular perception toolkit and Agentic Reinforcement Learning to internalize reasoning and tool‑use strategies, achieving competitive results on online video benchmarks.

By Siwei Wen, Zhangcheng Wang, Xingjian Zhang, Lei Huang, Wenjun Wu