arXiv AI

Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models

The paper examines how hallucinations arise in multi-stage video‑understanding agents by aligning existing benchmarks with the stages of temporal grounding, visual observation, and reasoning. It introduces a causal stage‑intervention protocol that isolates each stage while keeping the downstream task constant, revealing that grounding errors dominate downstream hallucinations and that correct region location matters more than precise temporal overlap. The study also shows that current benchmark scores poorly predict causal sensitivity and can fail under distribution shift, advocating for stage‑aware evaluation methods.

arXiv Computer Vision
Sep 18

Stress Tests REVEAL Fragile Temporal and Visual Grounding in Video-Language Models

The paper introduces REVEAL, a diagnostic benchmark that stresses Video‑Language Models (VidLMs) on five controlled probes—camera‑motion sensitivity, cross‑frame integration, video sycophancy, language‑only shortcuts, and temporal expectation bias—to assess how well these models encode and use visual evidence. Experiments on 12 VidLMs reveal systematic failures: some visual signals are never reliably encoded, while others are overridden by model priors, leading to performance below chance on several probes that humans solve with high accuracy. Mechanistic probes further pinpoint where and why visual evidence is lost, demonstrating that under assertive prompts a model’s output becomes nearly invariant to real versus random video input, rendering visual evidence causally inert.

By Sethuraman T V, Savya Khosla, Aditi Tiwari, Vidya Ganesh, Rakshana Jayaprakash, Aditya Jain, Vignesh Srinivasakumar, Onkar Kishor Susladkar, Srinidhi Sunkara, Aditya Shanmugham, Rakesh Vaideeswaran, Abbaas Alif Mohamed Nishar, Simon Jenni, Rohan Maheshwari, Derek Hoiem
arXiv Computer Vision
Aug 21

Video Evidence to Reasoning Efficient Video Understanding via Explicit Evidence Grounding

arXiv:2601. 07761v2 Announce Type: replace Abstract: Large Vision-Language Models (LVLMs) face a fundamental dilemma in video reasoning: they are caught between the prohibitive computational costs of verbose reasoning and the hallucination risks of efficient, ungrounded approaches.

By Yanxiang Huang, Guohua Gao, Zhaoyang Wei