TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.
By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
OVIBench introduces the first standardized benchmark for evaluating vision‑language models on Online Video Question Answering under Interruption, a realistic setting where users can interrupt the model during answer generation. The benchmark categorizes interruptions into Cancellation, False Trigger, and Correction, supports both open‑ended and multiple‑choice tasks, and provides an offline simulation protocol plus a multi‑dimensional metric suite. Experiments show that OVIBench can distinguish models’ interruption‑handling abilities, particularly in following correction requests, and that fine‑tuning on the newly created OVI‑Train dataset yields significant performance gains.
By Naiming Liu, Zhiheng Wu, Shuning Wang, Tie Zhang, Bowen Liu, Tong Wang
arXiv:2606. 26904v1 Announce Type: cross Abstract: Video reasoning language models implicitly assume that every input frame is equally reliable.
By Yangfan He, Yujin Choi, Jaehong Yoon
arXiv:2606. 24797v1 Announce Type: cross Abstract: Recent advances in Video Large Language Models (Video-LLMs) have yielded promising performance on video question answering (VideoQA).
By Linpeng Huang, Weixing Chen, Zexin Chen, Yang Liu, Liang Lin
arXiv:2607.01469v3 Announce Type: replace
Abstract: Agentic Video Question Answering (VideoQA) systems produce answers through adaptive reasoning and tool-use trajectories, yet standard practice eval...
By Rama AlHamidi, Aseel Mohamed, Rasul Khanbayov, Mohamed Rayan Barhdadi, Erchin Serpedin, Hasan Kurban
The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.
By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.
By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu
arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.
By Ali Alavi
arXiv:2603. 10652v3 Announce Type: replace-cross Abstract: In real-world deployment, vision-language models often encounter disturbances such as weather, occlusion, and camera motion.
By Yangfan He, Changgyu Boo, Jaehong Yoon
Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode? Benchmarks increasingly score this kind of task, and the strongest open-source 7--8B models now reach 37--38% on InfiniBench's global appearance task, which asks exactly that.
arXiv:2607. 11078v1 Announce Type: cross Abstract: Can a Video Large Language Model (Video-LLM) follow one person through a long video, keeping track of who they are well enough to report, in order, how their outfit changes across a full TV episode?
By Mohammad Al-Ratrout, Shayla Sharmin, Aditya Raikwar, Roghayeh Leila Barmaki
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee