arXiv AI

Event-Conditioned Diagnostics of Kinematic, Contact, and Object-Permanence Structure in Passive Object-State World Models

The paper introduces a diagnostic protocol to examine how passive object-state world models encode event-conditioned latent physical structure. It evaluates GRU, Transformer-lite, and RSSM-lite models on a dataset featuring free motion, collision, and occlusion events, finding that each architecture learns predictive dynamics and that event context shifts the emphasis among kinematic, contact, and object-permanence readouts. Functional sensitivity tests reveal contact-related structure during collisions and object-permanence structure during occlusions, supporting the presence of event-conditioned latent structure without explicit physical modules.

arXiv Computer Vision
4d ago

Foresight at the Event Boundary: Evaluating Physical Prediction in Video World Models

The paper introduces an event‑anchored evaluation protocol for video world models, using 62 free‑fall recordings and 124 clips with detailed release and impact annotations. It finds that while some models (Runway, Veo) can generate release and impact events with high accuracy, they often start them late, and others (Cosmos‑Predict‑2.5, MAGI‑1) rarely produce measurable consequences. A human study shows that people’s predictions align with recorded futures but also reveal ambiguity in plausible continuations, highlighting that physical foresight requires initiating, timing, and realizing motion correctly.

By Estela Monserrat Arriaga Santana (National Autonomous University of Mexico), Julian Rosas Scull (National Autonomous University of Mexico), Eh\'ecatl Sacamch'en N\'u\~nez Rico (National Autonomous University of Mexico), Hugo Jair Escalante (University of Texas at El Paso)
arXiv Machine Learning
Jul 30

What Can Latent World Models Know? Physical Parameter Identifiability in Multimodal Predictive Representations

arXiv:2607. 27017v1 Announce Type: new Abstract: A central premise of latent world models is that predicting the future forces a representation to internalize the physics of its environment.

By Kaizhen Tan (New York University, Carnegie Mellon University), Xin Xu (Carnegie Mellon University), Siru Tao (Carnegie Mellon University), Hanzhe Hong (Carnegie Mellon University), Yang Feng (Columbia University), Heqing Du (Columbia University)
arXiv AI
Aug 24

ForeTime-VLA: Causal Future-Token Distillation from a World Action Model for Conveyor-Belt Manipulation

ForeTime‑VLA is a causal vision‑language‑action policy that distills future‑aware representations from a frozen Fast‑WAM teacher, enabling it to anticipate contact events during conveyor‑belt manipulation. The method compresses current and future video latents into a 64‑dimensional target, uses an eight‑frame history encoder to predict this target along with manipulation phase and time‑to‑transition, and conditions a VLM prefix on future tokens and phase. On a deduplicated conveyor‑belt dataset, ForeTime‑VLA reduces test MAE by 2.63% and L2 by 3.02%, while real‑robot experiments show significantly higher grasp success rates compared to the next‑best reference. whyItMatters":"The approach demonstrates that distilling future‑token knowledge from a world‑action model can improve dynamic manipulation performance without the computational cost of running the teacher at inference time."

By Siyuan Ma, Yutian Zhang, Boshi Zhang, Qinglian Wu, Jiaqi Zhai, Dong Wei, Xiaojin Huang
arXiv Machine Learning
Aug 27

JEPA-x: Cross-Predictive Physics Grounding for Forecastable Latent Dynamics

JEPA-x is a cross‑predictive physics grounding method that aligns visual latent dynamics with privileged physical trajectories. By treating visual observations and physical states as two views of the same action‑conditioned trajectory and sharing a predictor, it forces the model to learn a common transition rule for both modalities. The physical branch is only used during training, so deployment incurs no extra cost, and the approach significantly reduces rollout drift and boosts control success across a multi‑task suite.

By Kehan Wen, Ziming Li, Siyuan Luo, Fan Shi