Stable Scores, Unstable Answers: Frame Phase and Option Order in Video Multiple-Choice Evaluation
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The paper investigates how changes in evaluation settings—such as camera angles and caption wording—affect the rankings of text-to-3D generators. Using 300 fixed scenes and varying eight render and caption factors, the authors find that configuration variance often exceeds generator variance, leading to frequent shifts in the top-scoring model across 19 alignment evaluators. They conclude that observed winner changes are descriptive rather than definitive, and recommend detailed reporting of generator, score, and protocol specifics to account for uncertainty.
arXiv:2509.01167v3 Announce Type: replace-cross Abstract: Vision-language models (VLMs) can ingest only a limited number of video frames, making frame selection a practical necessity. But do current...
A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions. 1.
arXiv:2607. 12304v1 Announce Type: cross Abstract: A score on a temporal video question answering benchmark is meant to measure that a model has temporal understanding, but it conflates two questions.
arXiv:2602.10639v2 Announce Type: replace Abstract: Video Large Language Models (VideoLLMs) have achieved strong performance on video understanding tasks, yet existing benchmarks evaluate only what m...
The study investigates how long‑video language models decide which frames to keep, compress, and reuse, testing each decision in isolation across six selection rules, three benchmarks, and two answering models. It finds that selecting frames based on queries yields the biggest performance boost, that halving spatial resolution costs little, and that reallocating saved tokens to more compressed frames can further improve accuracy. The work also highlights the importance of a unified evaluation harness to avoid misleading comparisons.