Why Do Conventional World Models Fail to Learn Cellular Automata?
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
arXiv:2609.06102v2 Announce Type: cross Abstract: Cellular automata is a local computation paradigm where complex behavior can arise from local interactions between simple functions. This paradigm ha...
arXiv:2608.30692v1 Announce Type: new Abstract: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world. W...
arXiv:2607. 27320v1 Announce Type: cross Abstract: Field-level inference of cosmological initial conditions from galaxy surveys requires a forward model that is simultaneously accurate in the non-linear regime, computationally efficient, and fully differentiable.
The paper introduces a sparse, residual world model that focuses on predicting only the changes in a scene by using a per-object change gate and a residual delta head. On a MuJoCo tabletop pushing benchmark, this approach outperforms a dense multilayer perceptron, achieving 2.5 to 4.6 times better next‑state pose accuracy with 8.6 to 11.1 times fewer parameters, maintaining high change‑detection F1 scores, and showing strong transfer across object counts. In autoregressive rollout and sampling‑based planning, the sparse model accumulates less error and enables successful planning where dense models fail.
arXiv:2603. 16689v2 Announce Type: replace Abstract: Next-token predictors often appear to develop internal representations of the latent world and its rules.
The paper investigates how transformers can possess a world model despite exhibiting behavioral failures. Using TaxiGPT, a transformer trained on random Manhattan walks, the authors show that the model internally represents intersections, streets, and its position, and uses a goal compass for navigation. They attribute failures to interference between overlapping intersection features and demonstrate that affordance packing mitigates these errors, concluding that world‑modeling abilities emerge at distinct training stages and should be studied mechanistically rather than merely observed behaviorally.