arXiv Machine Learning

Have I Seen Enough? Frozen Video-Language Models Encode Evidence Readiness

The paper investigates whether frozen video‑language models inherently encode a signal indicating whether sufficient evidence has been observed to answer a question. By training linear probes on seven byte‑identical models, the authors demonstrate that these models contain a readable evidence‑readiness signal with AUROC ranging from 0.733 to 0.905, even when the probe is trained without any footage from the benchmark family. The signal is question‑conditioned, remains robust when the model answers incorrectly, and outperforms traditional uncertainty estimators; it can be leveraged as a Readiness Gating policy that improves answer accuracy by up to 9.75 percentage points without extra computational cost.

arXiv AI
Sep 30

Watch-Think-Interact: Bootstrapping Long-Horizon Multi-Turn Streaming Video Reasoning with Reinforcement Learning

The paper introduces Watch-Think-Interact (WTI), a closed-loop framework for multi-question streaming video reasoning that maintains compact natural-language memory entries linked to video time ranges. WTI decides whether to answer, continue watching, or recall relevant past intervals for each question, avoiding replay of the full history. The authors build a large dataset, WTI-82K, and a training method, Stream-GDPO, achieving state‑of‑the‑art performance on StreamingBench and OVO-Bench.

By Ziheng Huang, Yicheng Bao, Xueheng Li, Zhenkun Gao, Bangwei Liu, Kunquan Li, Yuxiang Shen, Bangyan Li, Xuejiao Wang, Changbo Wang, Gaoqi He
arXiv Computation and Language
Sep 28

TRACE: Temporal Audit and Condition-aware Evaluation of Streaming Video Understanding

TRACE (Temporal Audit and Condition-aware Evaluation) is a new benchmark and evaluation framework for streaming video understanding that explicitly records when evidence becomes valid, how visual history is maintained, and how responses are triggered. It combines temporally audited visual tasks, evidence timing, instruction-dependent trigger annotations, a unified causal Core–Adapter protocol, and multidimensional reporting of answer quality, timeliness, response-selection behavior, workload, completion, and reliability. Using 1,240 records from 517 videos, TRACE evaluated eight publicly available models in eight configurations, revealing that similar QA accuracy can hide significant differences in completion, answer validity, generation workload, and proactive performance metrics such as response delay, false alarms, and missed target windows.

By Yibo Ma, Qianqian Zhang, Peng Liu, Tiancheng Zhao
arXiv AI
Jun 24

video-SALMONN-R$^3$: Learning to ReWatch, ReAsk, and ReAnswer for Efficient Video Understanding

arXiv:2606. 24477v1 Announce Type: cross Abstract: Video large language models (LLMs) are often constrained by computation and memory budgets, leading them to use reduced frame rates and spatial resolutions, which may cause them to miss critical information for question answering (QA).

By Yixuan Li, Guangzhi Sun, Yudong Yang, Wei Li, Zejun MA, Chao Zhang
arXiv Computer Vision
3d ago

Behavior Pack Optimization for Video MLLM Post-Training

The paper introduces Behavior Pack Optimization (BPO), a post‑training method for video multimodal large language models that replaces single‑response rewards with a set of outputs across counterfactual views. BPO enforces stability when interventions are irrelevant, sensitivity when key evidence is removed, and abstention when no evidence remains, using an anchor‑relative advantage to keep the objective stable with small pack sizes. Experiments on datasets such as TempCompass, MVBench, and NExT‑QA show that BPO improves macro accuracy and abstention metrics for models like Qwen2.5‑VL‑7B‑Instruct, with gains that transfer to other benchmarks and models.

By Zhaolu Kang, Shiyu Liu, Tailong Luo, Wei Zhang, Yingjie He, Lei Wei, Guansu Wang, Liang He, Siheng Wang, Guangyuan Dong, Jiaqi Su, Shuang Chen, Haoyu Ji, Qishi Zhan, Kaiyue Zhou
arXiv Machine Learning
Jun 2

Perception First: A Frontier Native-Video Model with Self-Consistency for Implicit Video Question Answering

arXiv:2606. 01485v1 Announce Type: cross Abstract: We describe our submission to the VRR Challenge @ CVPR 2026, built on the \emph{ImplicitQA} / \emph{VRR-QA} benchmark~\cite{implicitqa}: multiple-choice video question answering in which answers are deliberately \emph{not} observable in any single frame and must be inferred from spatial layout, motion, depth, viewpoint, causality, and social context across discontinuous frames of creative video.

By Ali Alavi
arXiv Computer Vision
Sep 28

STORM-Bench: Evaluating Online Video QA under Evolving and Incomplete Evidence

STORM-Bench is a new benchmark for online video question answering that evaluates models’ ability to track state transitions and selectively abstain when visual evidence is insufficient. It contains 5,736 questions across 630 short, change‑dense episodes in five egocentric domains and two simulation subsets, with questions stratified by change intensity and answerability. The benchmark introduces STORM‑BR, a harmonic metric that reveals abstention failures and overconfidence on uncertain queries, showing that traditional accuracy masks gaps in epistemic reliability and state tracking.

By Siru Zhong, Shenghan Tan, Rihong Yan, Xiaohui Lv, Yuzheng Zhuang, Shuai Tao, Wulong Liu, Haohuan Fu, Yuxuan Liang
arXiv Computer Vision
Sep 22

Look Where It Counts: A Free, Label-Free Visual Evidence Signal for Fine-Grained Vision-Language Reasoning

The paper introduces a free, label‑free visual evidence signal that improves fine‑grained vision‑language reasoning. By selecting image crops that maximize the model’s answer distribution peak, the method locates answer‑bearing regions without training or annotations, boosting accuracy from 70 % to 85 %. The evidence gap also complements model confidence, enabling better correctness prediction and error flagging.

By Santi Ram Tiwari, Nihal Naik, Devbrat Pandey, Nishant Sinha
arXiv Computation and Language
Aug 25

TRACE: Temporal Retrieval with Anchored and Convergent Evidence for Long-Horizon Video Understanding

The paper introduces TRACE, a training‑free agent for long‑video understanding that grounds answers in raw visual clips and builds an evidence bundle until the answer stabilises. It also presents VES‑Bench, a 600‑question benchmark over 348 public long videos that audits whether decoded frames cover all necessary evidence intervals at three strictness levels. TRACE achieves high accuracy on VES‑Bench (63.5% audit accuracy) and remains competitive on other video‑understanding benchmarks while using far fewer frames than uniform decoding.

By Pengyiang Liu, Junbo Niu, Xiaoyang Hu, Zhongyue Shi, Zitian Wang, Linjiang Huang, Si Liu