arXiv:2606. 20104v1 Announce Type: cross Abstract: Perception for action suggests that representations of the world should be shaped not by visual fidelity alone, but by their relevance for actions.
By Petr Ivashkov, Randall Balestriero, Bernhard Sch\"olkopf
arXiv:2603. 16689v2 Announce Type: replace Abstract: Next-token predictors often appear to develop internal representations of the latent world and its rules.
By Sasha Brenner, Thomas R. Kn\"osche, Nico Scherf
arXiv:2607. 22430v1 Announce Type: new Abstract: Learning world models that infer environment dynamics from high-dimensional observations and predict outcomes under candidate actions is central to planning and control.
By Xiangteng Zhang, Yang Guan, Bo Zhang, Ya-Qin Zhang, Shengbo Eben Li
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
Autonomous driving requires planning under both semantic constraints and predictive dynamics. Existing end-to-end driving approaches, however, typically emphasize only one side of this requirement: Vision-Language-Action (VLA) models exploit VLM priors for semantic reasoning, while World Action Models (WAMs) provide future-aware prediction through generative world modeling.
arXiv:2606. 09311v1 Announce Type: new Abstract: Joint Embedding Predictive Architectures (JEPAs) have shown promising world modeling capabilities, enabling planning in latent space by optimizing action trajectories using methods like the Cross-Entropy Method (CEM).
By Sergi Masip, Jonathan Swinnen, Yutong Hu, Renaud Detry, Tinne Tuytelaars