The paper introduces IMPLY, a method that evaluates world-model rollouts by inferring the physics implied by each rollout and scoring them based on how well a single object explains all rollouts, anchored to two calibration pushes the model has observed. Unlike traditional self-consistency checks that only compare futures against each other, IMPLY requires physical consistency with known pushes, revealing when a model incorrectly internalises an object. Experiments show that self-consistency alone can misidentify models, whereas anchored disagreement accurately distinguishes correct from incorrect object tracking and can select rollout sets nearly as well as an oracle with ground‑truth knowledge.
By Aman Mehta, Riya Baviskar
OneWorld introduces a shared‑mechanism counterfactual generation framework that jointly models multiple action‑conditioned futures using a common latent physical mechanism. By inferring distributions over latent mechanisms for each action‑outcome branch and aggregating them into shared‑world evidence, the model enforces consistency across interventions while preserving distinct action outcomes. Experiments in controlled environments demonstrate that OneWorld improves cross‑intervention physical consistency without sacrificing single‑rollout prediction quality.
By Ke He, Yichen Ding, Bin Yang
arXiv:2607. 00276v1 Announce Type: cross Abstract: Current large-language-model (LLM) physics benchmarks are usually scored by answer accuracy, which cannot distinguish genuine reasoning from recall of familiar problem patterns and reveals little about where a model's reasoning breaks down.
By Dong Zhang
The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.
By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
arXiv:2608. 05670v1 Announce Type: new Abstract: A model's agreement across perturbed inputs is used both as a label-free reliability signal and as a self-training target, on the premise that agreement tracks correctness.
By Rasul Khanbayov, Hasan Kurban
arXiv:2607. 11598v1 Announce Type: new Abstract: There are two standard ways to spend more compute at test time: let a model reason longer, or sample more attempts and keep one.
By Bojie Li, Noah Shi