arXiv Computer Vision

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

STORM-Bench is a new benchmark for online video question answering that evaluates models’ ability to track state transitions and selectively abstain when visual evidence is insufficient. It contains 5,736 questions across 630 short, change‑dense episodes in five egocentric domains and two simulation subsets, with questions stratified by change intensity and answerability. The benchmark introduces STORM‑BR, a harmonic metric that reveals abstention failures and overconfidence on uncertain queries, showing that traditional accuracy masks gaps in epistemic reliability and state tracking.

arXiv Computation and Language
6d ago

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.

By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv AI
Aug 25

OVIBench: Benchmarking Online Video Question Answering under Interruption

OVIBench introduces the first standardized benchmark for evaluating vision‑language models on Online Video Question Answering under Interruption, a realistic setting where users can interrupt the model during answer generation. The benchmark categorizes interruptions into Cancellation, False Trigger, and Correction, supports both open‑ended and multiple‑choice tasks, and provides an offline simulation protocol plus a multi‑dimensional metric suite. Experiments show that OVIBench can distinguish models’ interruption‑handling abilities, particularly in following correction requests, and that fine‑tuning on the newly created OVI‑Train dataset yields significant performance gains.

By Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang
arXiv AI
4d ago

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.

By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv Computation and Language
Aug 25

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.

By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu
arXiv Machine Learning
Jun 2

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.

By Ali Alavi
Hugging Face Trending Papers
Jul 13

Do Video-LLMs Actually Watch? Diagnosing Character-Tracking Failures in Long-Form Video

Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7--8B models now reach 37--38% on InfiniBench's global appearance task, which asks exactly that.