arXiv Computer Vision

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

arXiv AI
6d ago

OneWorld: Learning Consistent Physics Across Actions in World Models

OneWorld introduces a shared‑mechanism counterfactual generation framework that jointly models multiple action‑conditioned futures using a common latent physical mechanism. By inferring distributions over latent mechanisms for each action‑outcome branch and aggregating them into shared‑world evidence, the model enforces consistency across interventions while preserving distinct action outcomes. Experiments in controlled environments demonstrate that OneWorld improves cross‑intervention physical consistency without sacrificing single‑rollout prediction quality.

By Ke He, Yichen Ding, Bin Yang
arXiv AI
Sep 15

One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

The paper investigates how multi‑modal world models can produce inconsistent outputs across different modalities, such as a video showing a ball not rebounding while a text description indicates it should. It defines two types of misalignment—internal (between modalities) and external (against a physical environment)—and introduces a physics‑grounded pipeline to measure these discrepancies. Experiments across multiple settings reveal that while the model’s language output matches the true environment, its video output frequently disagrees, indicating current unified backbones struggle with simultaneous reasoning, consistency, and physical fidelity.

By Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe, Manish Bhattarai
arXiv Computer Vision
Sep 10

MotionBlind: Probing the Illusion of Motion Understanding in Video-LLMs

arXiv:2609.09528v1 Announce Type: new Abstract: Video large language models (Video-LLMs) are increasingly used as the perceptual front end of world models, a role that assumes they can read motion: h...

By Dhairya Bhatia, Bishoy Galoaa, Oliver Fritsche, Shahid Kamal, Muhammad Obaidullah Abdul Salam, Umer Saleem, Om Rastogi, Frania Felix Chettiar, Nesli Erdogmus, Sarah Ostadabbas
arXiv AI
Aug 28

PAWBench: How Far Are We from Probabilistically Aligned World Modeling?

PAWBench introduces a benchmark to evaluate whether video generation models can act as probabilistically aligned world models, meaning they should reproduce not just plausible trajectories but the full distribution of possible behaviors from the same initial conditions. The authors formalize probabilistic alignment as a distributional criterion and provide PAWEval, an outcome-level protocol that turns repeated video rollouts into empirical distributions over physical behaviors. Across 50 scenarios and eleven current systems, none consistently matched reference probabilities or captured the full range of valid behaviors, highlighting a significant gap in current video generators.

By Yuandong Pu, Le Zhuo, Sayak Paul, Gabriel Jorge Menezes, Avram {\DJ}or{\dj}evi\'c, Shiyang Li, Yifan Zhou, Bin Fu, Wenlong Zhang, Junjun He, Yu Qiao, Yihao Liu, Jingbo Xing, Xi Chen
arXiv Computer Vision
4d ago

ReWorld-Track: A Recursive Event World Model for Language-Guided Multi-Camera Tracking

ReWorld-Track introduces a recursive event world model for language‑guided multi‑camera tracking that explicitly carries association uncertainty into future predictions. By treating candidate matches and waiting as alternative target states, the model updates a persistent recurrent belief that preserves uncertainty across successive observations. This approach improves identity continuity and next‑camera accuracy, achieving HOTA scores of 65.19 on CityFlowV2 and 45.36 on MTMMC, and reducing median arrival‑time error from 0.78 s to 0.71 s.

By Haoyang Wu, Shoudong Han, Chaoyue Li, Sijia Chen, Zhenyang Xie, Wang sihan