TempCore: Are Video QA Benchmarks Temporally Grounded?
Read the original on arXiv Machine Learning →The Flow has not summarised this story yet — read it at arXiv Machine Learning.
The Flow has not summarised this story yet — read it at arXiv Machine Learning.
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
The paper introduces REVEAL, a diagnostic benchmark that stresses Video‑Language Models (VidLMs) on five controlled probes—camera‑motion sensitivity, cross‑frame integration, video sycophancy, language‑only shortcuts, and temporal expectation bias—to assess how well these models encode and use visual evidence. Experiments on 12 VidLMs reveal systematic failures: some visual signals are never reliably encoded, while others are overridden by model priors, leading to performance below chance on several probes that humans solve with high accuracy. Mechanistic probes further pinpoint where and why visual evidence is lost, demonstrating that under assertive prompts a model’s output becomes nearly invariant to real versus random video input, rendering visual evidence causally inert.
arXiv:2608.29958v1 Announce Type: new Abstract: Long videos contain far more visual content than Large Vision-Language Models (LVLMs) can process under a fixed visual-token budget, making frame selec...
arXiv:2602.10639v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what m...
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.
The paper introduces VT-Contrast, a representation-level temporal counterfactual objective designed to improve temporal understanding in Video Language Models (VideoLMs). By supervising late-layer last-frame video tokens and contrasting order-preserving views with reordered counterfactuals graded by Kendall tau distance, VT-Contrast addresses the mismatch between ordered video input and text-based supervision. The method requires no architectural changes, is compatible with various VideoLM training tasks, and demonstrates improved performance on temporal understanding benchmarks.