Beneath the Scores: Rethinking Hallucination Evaluation for Video Understanding Models
Read the original on arXiv AI →The paper examines how hallucinations arise in multi-stage video‑understanding agents by aligning existing benchmarks with the stages of temporal grounding, visual observation, and reasoning. It introduces a causal stage‑intervention protocol that isolates each stage while keeping the downstream task constant, revealing that grounding errors dominate downstream hallucinations and that correct region location matters more than precise temporal overlap. The study also shows that current benchmark scores poorly predict causal sensitivity and can fail under distribution shift, advocating for stage‑aware evaluation methods.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.