arXiv AI

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

arXiv:2608. 16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested.

arXiv Computer Vision
Sep 22

Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning

Video foundation models now match human accuracy on physical‑reasoning benchmarks, but a new distributional evaluation framework shows that their predictions diverge markedly from human judgments. On the Physion benchmark, ViT‑L models such as V‑JEPA2, VideoMAE‑v2, and DINOv2 achieve near‑human accuracy yet exhibit a 26.4% model‑human disagreement, far above the 4.8% human‑human disagreement, and lower agreement (kappa ~0.48 vs. 0.91). The divergence varies by task: models excel at geometric reasoning but lag on gravitational dynamics and causal chains, indicating they rely on statistical regularities rather than explicit forward simulation.

By Fanhong Li, Shurui Zheng, Zi Yin, Junbo Cui, Lei Ji, Jia Liu
arXiv AI
Aug 28

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

PAWBench introduces a benchmark to evaluate whether video generation models can act as probabilistically aligned world models, meaning they should reproduce not just plausible trajectories but the full distribution of possible behaviors from the same initial conditions. The authors formalize probabilistic alignment as a distributional criterion and provide PAWEval, an outcome-level protocol that turns repeated video rollouts into empirical distributions over physical behaviors. Across 50 scenarios and eleven current systems, none consistently matched reference probabilities or captured the full range of valid behaviors, highlighting a significant gap in current video generators.

By Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi\'c, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
arXiv Machine Learning
Sep 11

New Evidence, Same Choice: Testing Physical Experiment Selection in Vision Language Models

The paper introduces a new benchmark for vision‑language models that tests their ability to decide whether to answer a physics question immediately or to request additional experimental evidence. Each problem presents one measurement image and four possible physical worlds defined by two masses and two values of another property; the model must either stop and answer or choose the cheapest experiment that resolves the question. Across six open models and 144 parameter sets, the models almost always repeat the same action even when the optimal choice changes, and only a single model gets both decisions correct on 5.9% of cases.

By Sourajit Saha, Shubhashis Roy Dipta, Nobin Sarwar, Shaswati Saha, Yuxuan Jiang, Siyuan Li, Qiheng Wang
arXiv Computer Vision
Aug 28

R2M-Bench: Evaluating Revisit Memory via Relative Consistency in Interactive Video World Models

R2M-Bench is a benchmark that evaluates revisit memory in interactive video world models by comparing a revisit pair to two control pairs from the same rollout: a gap‑matched non‑revisit pair and a short‑range pair. It introduces MemoryGain (MG) and Normalized Memory Ratio (NMR) to quantify the revisit advantage over generic temporal stability and normalize it by short‑to‑baseline dynamics. Across 300 instances and seven models, NMR correlates with human judgments and reduces the influence of slow‑motion artifacts, with DreamX‑World‑Memo achieving the highest NMR.

By Qiwen Gu, Bingjie Gao, Rui Chen, Geng Li, Jifan Li, Qishuai Wen, Li Niu, Jing Tang, Xiangxiang Chu, Junqiao Zhao
arXiv Computer Vision
4d ago

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
arXiv AI
Sep 10

CALIPER: Clean Scenes Cannot Rank Physical Inference in Pretrained Visual Representations

CALIPER is a new benchmark that tests whether pretrained visual encoders can infer physical properties such as mass and friction from images. The test involves striking an object twice at known speeds, showing a third strike only up to contact, and asking a linear readout on frozen features to predict how far the object slides. Results show that in clean, fixed‑camera scenes all representations perform similarly, but when camera, lighting, and clutter are varied, only encoders that truly infer physics—like V‑JEPA 2—maintain performance, while random or raw pixel representations fail.

By Aman Mehta, Riya Baviskar