arXiv Computer Vision By Fanhong Li, Shurui Zheng, Zi Yin, Junbo Cui, Lei Ji, Jia Liu

Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning

Read the original on arXiv Computer Vision →

Video foundation models now match human accuracy on physical‑reasoning benchmarks, but a new distributional evaluation framework shows that their predictions diverge markedly from human judgments. On the Physion benchmark, ViT‑L models such as V‑JEPA2, VideoMAE‑v2, and DINOv2 achieve near‑human accuracy yet exhibit a 26.4% model‑human disagreement, far above the 4.8% human‑human disagreement, and lower agreement (kappa ~0.48 vs. 0.91). The divergence varies by task: models excel at geometric reasoning but lag on gravitational dynamics and causal chains, indicating they rely on statistical regularities rather than explicit forward simulation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computer Vision.

arXiv Computer Vision
2d ago

SYNCR: A Cross-Video Reasoning Benchmark with Synthetic Grounding

SYNCR is a synthetic benchmark designed to evaluate multimodal large language models on cross‑video reasoning. It contains 4,000 question‑answer pairs across 4,827 unique videos, covering tasks in temporal alignment, spatial tracking, comparative reasoning, and holistic synthesis. The benchmark reveals a significant performance gap between current models and humans, with models excelling at temporal ordering but struggling with precise physical and spatial reasoning.

By Sara Ghazanfari, Siddharth Garg, Prashanth Krishnamurthy, Farshad Khorrami
arXiv Computation and Language
Sep 1

World Models Meet Language Models: On the Complementarity of Concrete and Abstract Reasoning

The paper introduces a framework that combines world models, which generate concrete visual rollouts of possible futures, with multimodal large language models (MLLMs) that perform abstract reasoning. It proposes a controlled concrete reasoning approach and a new training method called Privileged‑Future On‑Policy Self‑Distillation (PF‑OPSD), which uses ground‑truth future videos as privileged teacher context during training while the student model never sees true futures at test time. Experiments on two human‑verified benchmarks, VRQABench and OpenWorldQA, show that PF‑OPSD improves performance by about 10–11% over baselines and enhances robustness to noisy or conflicting rollouts.

By Yucheng Zhou, Wei Tao, Yiwen Guo, Jianbing Shen
arXiv AI
Aug 18

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

arXiv:2608. 16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested.

By Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke