arXiv Computation and Language

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.

arXiv Computation and Language
Aug 25

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.

By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu
arXiv Computer Vision
6d ago

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

STORM-Bench is a new benchmark for online video question answering that evaluates models’ ability to track state transitions and selectively abstain when visual evidence is insufficient. It contains 5,736 questions across 630 short, change‑dense episodes in five egocentric domains and two simulation subsets, with questions stratified by change intensity and answerability. The benchmark introduces STORM‑BR, a harmonic metric that reveals abstention failures and overconfidence on uncertain queries, showing that traditional accuracy masks gaps in epistemic reliability and state tracking.

By Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
arXiv AI
Sep 18

AVTrace: Diagnosing Audio-Visual Temporal Reasoning in Omni Models

AVTrace is a diagnostic suite designed to evaluate audio‑visual temporal reasoning in omni models. It covers tasks such as onset and span grounding, synchronization, next‑step prediction, cross‑modal localization, chain parsing, and event‑conditioned comprehension, providing 34,114 training examples and balanced development and test splits. Five open omni models were tested, all scoring below the majority‑label baseline on synchronization verification and showing low performance on chain parsing and event‑conditioned tasks, while parameter‑efficient temporal post‑training improved some metrics.

By Longyin Zhang, Parth Sakhare Mahendra, Chengwei Wei, Ning Zhang, Lim Ming Chong, Sirui He, Ai Ti Aw
arXiv AI
4d ago

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.

By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv Machine Learning
Sep 14

ProactiveBench: Can Streaming Video Models Really Interact Like Humans?

ProactiveBench evaluates streaming video models on their ability to interact proactively, rather than reactively. It tests models at one‑second intervals without explicit cues, using six subtasks that vary trigger ambiguity, timing tolerance, and response patterns. The benchmark measures both response and silence rates, distinguishing early, in‑window, and missed responses, and penalizes omissions and repetitions.

By Kaixuan Du, Xin Wan, YuKun Wang, Hang Zhang, Meng Cao, Dai Guan, Ming Chen, Ni Li
arXiv Computer Vision
2d ago

OneStreamer: Unifying Perception, Memory, and Proactive Response in Streaming Video Interaction

OneStreamer is a streaming video model that jointly learns to record evidence and respond to tasks through a shared proactive generation process. Its Proactive Hierarchical Caption Memory creates time‑grounded local‑detail captions and event summaries, while Proactive State Transition Learning reduces waiting states by supervising all output anchors. The authors also built a large OneStreamer‑1M dataset and show that a 4B model outperforms baselines on eight streaming video benchmarks, with ablations confirming the benefits of generated captions and PSTL.

By Xiangyu Zeng, Yuandong Yang, Zhiqiu Zhang, Yuhan Zhu, Xinhao Li, Qingyi Si, Dingyu Yao, Changlian Ma, Haoran Chen, Xinyu Chen, Yansong Shi, Junhao Zhou, Yifei Li, Jun Zhang, Chuanyu Qin, Chenxu Yang, Xinlei Yu, Kun Ouyang, Yuchen Shao, Qianshan Wei, Changhai Zhou, Jun Gao, Jiaqi Wang, Limin Wang
arXiv AI
Aug 20

Event-Causal RAG: A Retrieval-Augmented Generation Framework for Long Video Reasoning in Complex Scenarios

Event-Causal RAG (EC‑RAG) is a lightweight retrieval‑augmented framework designed for reasoning over ultra‑long and streaming videos. It segments video streams into semantically complete events using a dual visual‑audio sentinel mechanism, representing each event as a State‑Event‑State (SES) structure that captures pre‑event, event, and post‑event states. During question answering, bidirectional graph retrieval accesses relevant predecessor and successor events from a dual vector‑graph memory, and answers are generated using both this structured memory and the corresponding video evidence. The authors also introduce ECV‑1H, an hour‑scale long‑video QA benchmark with over 150 hours of untrimmed video and 1,251 human‑annotated QA pairs, where EC‑RAG achieves significant accuracy gains across multiple video foundation models while maintaining efficient streaming memory usage on a single RTX 5090 GPU.

By Peizheng Yan, Yu Zhao, Liang Xie, Juntong Qi, Mingming Wang, Erwei Yin