arXiv AI

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Fields in Passive Object-State World Models

arXiv:2606. 28455v1 Announce Type: cross Abstract: World models can predict future physical states, but prediction accuracy alone does not explain how physical information is organized and used inside their latent dynamics.

arXiv AI
Aug 25

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Structure in Passive Object-State World Models

The paper introduces a diagnostic protocol to examine how passive object-state world models encode event-conditioned latent physical structure. It evaluates GRU, Transformer-lite, and RSSM-lite models on a dataset featuring free motion, collision, and occlusion events, finding that each architecture learns predictive dynamics and that event context shifts the emphasis among kinematic, contact, and object-permanence readouts. Functional sensitivity tests reveal contact-related structure during collisions and object-permanence structure during occlusions, supporting the presence of event-conditioned latent structure without explicit physical modules.

By Yang Liu, Yuming Chen
arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi
arXiv Computer Vision
4d ago

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
arXiv AI
Sep 25

Underwater C3-JEPA: An Object-Centric Cross-View World Model for ROV Salvage

Underwater C³-JEPA is an object‑centric, cross‑view predictive world model designed for near‑field heavy‑load ROV salvage. It encodes synchronized multi‑camera RGB observations and vehicle control signals into task‑object and context tokens, fuses cross‑camera evidence via held‑out‑view attention, and predicts future states conditioned on control without using contact sensors. The model demonstrates superior transfer of task‑relevant information to downstream probes compared to a reconstruction‑free baseline, supports model‑predictive control and imagined‑rollout training, and validates its effectiveness on real underwater video by accurately recovering withheld camera states and maintaining predictive lead over persistence.

By Yuncong Yang, Jinlong Li, Yulong Xue, Feng Wu, Chunwen Zhang, Lei Qiao, Xuyang Wang
arXiv Machine Learning
Jul 30

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

arXiv:2607. 27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment.

By Kaizhen Tan (New York University, Carnegie Mellon University), Xin Xu (Carnegie Mellon University), Siru Tao (Carnegie Mellon University), Hanzhe Hong (Carnegie Mellon University), Yang Feng (Columbia University), Heqing Du (Columbia University)