arXiv AI By Geigh Zollicoffer, Minh Vu, Rajiv Ranasinghe, Manish Bhattarai

One Model, Two Physical Stories: Auditing Misalignment in Multi-Modal World Modeling

Read the original on arXiv AI →

The paper investigates how multi‑modal world models can produce inconsistent outputs across different modalities, such as a video showing a ball not rebounding while a text description indicates it should. It defines two types of misalignment—internal (between modalities) and external (against a physical environment)—and introduces a physics‑grounded pipeline to measure these discrepancies. Experiments across multiple settings reveal that while the model’s language output matches the true environment, its video output frequently disagrees, indicating current unified backbones struggle with simultaneous reasoning, consistency, and physical fidelity.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Machine Learning
Sep 14

IMPLY: Physically Anchored Consistency for World-Model Rollouts

The paper introduces IMPLY, a method that evaluates world-model rollouts by inferring the physics implied by each rollout and scoring them based on how well a single object explains all rollouts, anchored to two calibration pushes the model has observed. Unlike traditional self-consistency checks that only compare futures against each other, IMPLY requires physical consistency with known pushes, revealing when a model incorrectly internalises an object. Experiments show that self-consistency alone can misidentify models, whereas anchored disagreement accurately distinguishes correct from incorrect object tracking and can select rollout sets nearly as well as an oracle with ground‑truth knowledge.

By Aman Mehta, Riya Baviskar
arXiv AI
6d ago

OneWorld: Learning Consistent Physics Across Actions in World Models

OneWorld introduces a shared‑mechanism counterfactual generation framework that jointly models multiple action‑conditioned futures using a common latent physical mechanism. By inferring distributions over latent mechanisms for each action‑outcome branch and aggregating them into shared‑world evidence, the model enforces consistency across interventions while preserving distinct action outcomes. Experiments in controlled environments demonstrate that OneWorld improves cross‑intervention physical consistency without sacrificing single‑rollout prediction quality.

By Ke He, Yichen Ding, Bin Yang
arXiv Computer Vision
4d ago

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)