VideoSTF: Stress-Testing Output Repetition in Video Large Language Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper introduces REVEAL, a diagnostic benchmark that stresses Video‑Language Models (VidLMs) on five controlled probes—camera‑motion sensitivity, cross‑frame integration, video sycophancy, language‑only shortcuts, and temporal expectation bias—to assess how well these models encode and use visual evidence. Experiments on 12 VidLMs reveal systematic failures: some visual signals are never reliably encoded, while others are overridden by model priors, leading to performance below chance on several probes that humans solve with high accuracy. Mechanistic probes further pinpoint where and why visual evidence is lost, demonstrating that under assertive prompts a model’s output becomes nearly invariant to real versus random video input, rendering visual evidence causally inert.
arXiv:2607. 13305v1 Announce Type: cross Abstract: Benchmark accuracy in video large language models (LLMs) is often treated as evidence of visual understanding.
arXiv:2509.01167v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
arXiv:2609.39563v1 Announce Type: new Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes betwee...
arXiv:2608. 06361v1 Announce Type: new Abstract: Real-world video benchmarks provide broad coverage, but their fixed clips entangle event count, rate, duration, and visual complexity, making failure modes hard to isolate.
arXiv:2606. 09803v1 Announce Type: cross Abstract: We present \textbf{Echo-Memory}, a controlled study of memory mechanisms in action-conditioned world models.