arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
By Jae Joong Lee
The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.
By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.
By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).
By Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li, Zejun MA, Chao Zhang
arXiv:2606. 16353v1 Announce Type: cross Abstract: Streaming video understanding models must answer queries at any moment during an ongoing stream, using only what they have observed so far and under fixed memory and computation budgets.
By Haonan Ge, Yiwei Wang, Hang Wu, Yujun Cai
The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.
By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou
arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.
By Ali Alavi
STORM-Bench is a new benchmark for online video question answering that evaluates models’ ability to track state transitions and selectively abstain when visual evidence is insufficient. It contains 5,736 questions across 630 short, change‑dense episodes in five egocentric domains and two simulation subsets, with questions stratified by change intensity and answerability. The benchmark introduces STORM‑BR, a harmonic metric that reveals abstention failures and overconfidence on uncertain queries, showing that traditional accuracy masks gaps in epistemic reliability and state tracking.
By Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
arXiv:2609.09985v1 Announce Type: new
Abstract: Real-world video understanding requires integrating visual, audio, textual, and temporal evidence distributed across a video. Yet many pipelines use a...
By Sheng Li, Peng Liu, Qianqian Zhang, Tiancheng Zhao
The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.
By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
arXiv:2606. 06991v1 Announce Type: cross Abstract: Online Video Large Language Models (Video-LLMs) have advanced toward seamless human-AI interaction through frame-by-frame processing and proactive responding.
By Zhenyu Yang, Kairui Zhang, Shengsheng Qian, Weiming Dong, Changsheng Xu
The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.
By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu