arXiv AI By Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi\'c, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

Read the original on arXiv AI →

PAWBench introduces a benchmark to evaluate whether video generation models can act as probabilistically aligned world models, meaning they should reproduce not just plausible trajectories but the full distribution of possible behaviors from the same initial conditions. The authors formalize probabilistic alignment as a distributional criterion and provide PAWEval, an outcome-level protocol that turns repeated video rollouts into empirical distributions over physical behaviors. Across 50 scenarios and eleven current systems, none consistently matched reference probabilities or captured the full range of valid behaviors, highlighting a significant gap in current video generators.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

Video Generation Models: A Survey of Post-Training and Alignment

arXiv:2610.00812v1 Announce Type: cross Abstract: Video generation has rapidly progressed from short, low-quality clips to high-resolution, long-duration sequences with complex spatiotemporal dynamic...

By Chaoyu Li, Xiaoyi Gu, Yogesh Kulkarni, Eun Woo Im, Mohammadmahdi Honarmand, Zeyu Wang, Juntong Song, Fei Du, Xilin Jiang, Kexin Zheng, Tianzhi Li, Fei Tao, Pooyan Fazli
arXiv Computer Vision
Sep 3

SolarWM: Open Data and Scalable Training for Long-Horizon Video World Models

SolarWM is an open foundation for building interactive video world models, offering a reconfigurable multi‑source data engine that unifies 1.43 million clips from 10 datasets into a consistent, frame‑aligned format. It provides a backbone‑native adaptation framework that preserves native representations of models ranging from 5 B to 33 B parameters, and a three‑stage training recipe combining bidirectional adaptation, teacher‑forced autoregressive initialization, and distribution‑matching distillation. The resulting causal models can interact in real‑time over rollouts from minutes to hours, trained only on 5‑second sequences, and the project releases data, pipeline, recipes, weights, and framework for reproducible research.

By Junchao Huang, Guian Fang, Shengju Qian, Xianghao Kong, Zhuoran Zhao, Wei Huang, Yihua Du, Zixin Zhang, Justin Cui, Yuchao Gu, Yukang Chen, Xinting Hu, Tianyu He, Shaoshuai Shi, Zhuotao Tian, Xin Wang, Mike Zheng Shou, Li Jiang
arXiv AI
Sep 10

WorldAgen: Unified State-Action Prediction with Test-Time World Model Training

WorldAgen is a unified framework that jointly learns world modeling and action prediction using a shared Transformer backbone with two specialized heads. It introduces a Mixed Unidirectional Attention Mask to separate the world model and agent model, and enables Test-Time Training (TTT) by sampling exploratory actions and updating the world model with real state transitions. Experiments on CALVIN and LIBERO show that WorldAgen matches or surpasses state‑of‑the‑art methods, especially when TTT is applied to a few samples.

By Chi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang, Manling Li
arXiv AI
Aug 18

CaliBench: Are the Stochastic Dynamics of Video World Models Physically Calibrated?

arXiv:2608. 16829v1 Announce Type: cross Abstract: Video world models approximate the stochastic distribution of physical outcomes through generative sampling, but existing benchmarks score individual generations or compare distributions coarsely over a whole dataset, leaving the fine-grained aleatoric uncertainty of specific phenomena untested.

By Jonathan Sadeghi, Jenny Seidenschwarz, Jesse Allardice, Sirish Srinivasan, Benjamin Graham, Jeffrey Hawke
arXiv Computer Vision
Sep 22

Human-Level Accuracy, Non-Human Strategies: Revealing Model-Human Divergence in Video Physical Reasoning

Video foundation models now match human accuracy on physical‑reasoning benchmarks, but a new distributional evaluation framework shows that their predictions diverge markedly from human judgments. On the Physion benchmark, ViT‑L models such as V‑JEPA2, VideoMAE‑v2, and DINOv2 achieve near‑human accuracy yet exhibit a 26.4% model‑human disagreement, far above the 4.8% human‑human disagreement, and lower agreement (kappa ~0.48 vs. 0.91). The divergence varies by task: models excel at geometric reasoning but lag on gravitational dynamics and causal chains, indicating they rely on statistical regularities rather than explicit forward simulation.

By Fanhong Li, Shurui Zheng, Zi Yin, Junbo Cui, Lei Ji, Jia Liu