arXiv AI

Low-Rank Dynamics-Effective Latent Carriers for Counterfactual Rollout in Learned World Models

The paper investigates whether a small, directly addressable change in the hidden state of a learned world model can steer its future predictions along a desired counterfactual trajectory. Using a 192‑dimensional recurrent model in a two‑object collision setting, the authors identify low‑rank latent carriers—specifically a rank‑4 patch—that, when applied, successfully redirect a 12‑step autonomous rollout without further intervention. The study demonstrates that this compact intervention interface consistently works across independently trained checkpoints and intervention times, while various control experiments confirm the specificity of the effect.

arXiv AI
Sep 25

AD-WM: Action-Discriminative World Models for Counterfactual Model Predictive Control

AD-WM is a new action‑discriminative joint‑embedding world model designed for counterfactual model predictive control. It augments residual latent dynamics with action‑recovery regularization based on inverse dynamics and conditional mutual information, while discarding auxiliary heads at test time so that MPC remains unchanged. Experiments on OGBench‑Cube and other simulation environments show substantial gains in hard‑start success and mean success, and zero‑shot transfer to a Franka robot improves pick‑and‑place success from 42.2% to 71.1%.

By Jiabin Qiu, Zixuan Chen, Hongye Cao, Jieqi Shi, Jing Huo, Yang Gao
arXiv Machine Learning
Sep 1

The Intervention Gap in Latent World Models

The paper introduces the concept of intervention fidelity in latent world models, measuring whether a model’s open‑loop transitions align with actual environment interventions. Experiments on TD‑MPC2, Cheetah, and DreamerV3 show that high reward fit does not guarantee fidelity, and that self‑supervised models can outperform task‑anchored ones in preserving intervention effects. The authors propose a capture‑gated audit to localize failures and argue that fidelity must be directly audited on the model’s native interface.

By Donna Vakalis
arXiv AI
6d ago

OneWorld: Learning Consistent Physics Across Actions in World Models

OneWorld introduces a shared‑mechanism counterfactual generation framework that jointly models multiple action‑conditioned futures using a common latent physical mechanism. By inferring distributions over latent mechanisms for each action‑outcome branch and aggregating them into shared‑world evidence, the model enforces consistency across interventions while preserving distinct action outcomes. Experiments in controlled environments demonstrate that OneWorld improves cross‑intervention physical consistency without sacrificing single‑rollout prediction quality.

By Ke He, Yichen Ding, Bin Yang
arXiv Computer Vision
4d ago

CST-WM: A Causally Structured World Model for Embodied Visual Tracking

CST‑WM is a causally structured world model designed for embodied visual tracking, where a robot must keep a moving target visible and recover it after occlusion or drift. The model separates state into target‑evidence, robot, and observation branches, removing direct action‑to‑target‑evidence edges to prevent causal hallucination and instead letting actions influence evidence through robot motion and resulting views. Evaluated on EVT‑Bench, Habitat 3.0, and real‑world trials with a Unitree Go2 quadruped, CST‑WM outperforms reactive trackers and other world‑model baselines in following, distance control, safety, and re‑acquisition, achieving 20 of 30 successful real‑world recoveries versus 14 for TrackVLA.

By Junyi Hu, Shuaihang Yuan, Jiazhao Liang, Yi Fang
arXiv Machine Learning
Jul 14

A Control Theory of Predictability in Latent World Models

arXiv:2607. 10362v1 Announce Type: new Abstract: Latent world models are trained to predict future states in a learned representation and are then deployed inside a planner that selects actions by simulating them forward.

By Hanzhe You, Yonggang Zhang, Maohao Ran, Zhiqin Yang, Zhenyuan Zhang, Wei Xue, Jun Song, Xinmei Tian, Yike Guo
arXiv AI
Sep 3

Modeling What Changes: Sparse, Residual World Models for Object-Centric Manipulation

The paper introduces a sparse, residual world model that focuses on predicting only the changes in a scene by using a per-object change gate and a residual delta head. On a MuJoCo tabletop pushing benchmark, this approach outperforms a dense multilayer perceptron, achieving 2.5 to 4.6 times better next‑state pose accuracy with 8.6 to 11.1 times fewer parameters, maintaining high change‑detection F1 scores, and showing strong transfer across object counts. In autoregressive rollout and sampling‑based planning, the sparse model accumulates less error and enables successful planning where dense models fail.

By Param Thakkar, Parsika Paresh Shah, Manisha Sushant Gote