arXiv AI

One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

The paper investigates how multi‑modal world models can produce inconsistent outputs across different modalities, such as a video showing a ball not rebounding while a text description indicates it should. It defines two types of misalignment—internal (between modalities) and external (against a physical environment)—and introduces a physics‑grounded pipeline to measure these discrepancies. Experiments across multiple settings reveal that while the model’s language output matches the true environment, its video output frequently disagrees, indicating current unified backbones struggle with simultaneous reasoning, consistency, and physical fidelity.

arXiv Machine Learning
Sep 14

IMPLY: Physically Anchored Consistency for World-Model Rollouts

The paper introduces IMPLY, a method that evaluates world-model rollouts by inferring the physics implied by each rollout and scoring them based on how well a single object explains all rollouts, anchored to two calibration pushes the model has observed. Unlike traditional self-consistency checks that only compare futures against each other, IMPLY requires physical consistency with known pushes, revealing when a model incorrectly internalises an object. Experiments show that self-consistency alone can misidentify models, whereas anchored disagreement accurately distinguishes correct from incorrect object tracking and can select rollout sets nearly as well as an oracle with ground‑truth knowledge.

By Aman Mehta, Riya Baviskar
arXiv AI
6d ago

OneWorld: Learning Consistent Physics Across Actions in World Models

OneWorld introduces a shared‑mechanism counterfactual generation framework that jointly models multiple action‑conditioned futures using a common latent physical mechanism. By inferring distributions over latent mechanisms for each action‑outcome branch and aggregating them into shared‑world evidence, the model enforces consistency across interventions while preserving distinct action outcomes. Experiments in controlled environments demonstrate that OneWorld improves cross‑intervention physical consistency without sacrificing single‑rollout prediction quality.

By Ke He, Yichen Ding, Bin Yang
arXiv Computer Vision
4d ago

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
arXiv Machine Learning
1d ago

Stable and Counterfactually Robust Physical World Models from Imposed Structure and Learned Physics

The paper introduces a world model that learns to predict the evolution of physical systems while respecting key physical principles. By hard‑coding a general structure—generating dynamics from the gradient of a learned energy via a fixed reversible operator and imposing constraints on energy, dissipation, and interventions—the model achieves second‑law compatible dissipation, accurate responses to parameter changes, long‑term stability, and robustness to disturbances. Experiments on an electromagnetic cavity, a particle‑in‑cell grid, and shallow‑water fluid demonstrate that the model can recover accurate constitutive functions, distinguish conserving from dissipating regimes, and transfer learned physics to unseen conditions, outperforming unconstrained models.

By Yufeng Wang, Parivesh Priye, Lu Wei, Haibin Ling
arXiv AI
Aug 19

Where a New Concept Must Enter: Entry Point Gates Cross-Task Usability in Unified Multimodal Models

The paper investigates how new concepts can be integrated into unified multimodal models (UMMs) by separating generation and understanding objectives through a novel visual entity bound to a single task direction. Experiments show that the effectiveness of cross‑task usability depends on where the concept is injected into the shared computation, with a mid‑stack alignment objective achieving high concept acquisition with minimal loss to overall performance. The study highlights that unified weights alone are insufficient; the two directions must share a semantic format at the entry point for efficient concept integration.

By Zongyang Qiu, Yihan Wu, Kaixuan Fan, Bo Li, Hui Xiong
arXiv Machine Learning
Aug 21

Beyond Multimodal Alignment: Certifying Physical Language through Response Substitution and Ordered Execution

arXiv:2608. 19492v1 Announce Type: new Abstract: World models increasingly treat compact multimodal representations as interfaces between perception and physical interaction, yet existing probes do not establish whether different sensors carry the same executable meaning or whether that meaning survives a new action composition.

By Kaizhen Tan, Xin Xu, Siru Tao, Yixiao Li, Hanzhe Hong, Yang Feng, Heqing Du