The paper investigates the discrepancy between strong agents and their underlying weak world models in Atari Pong by reproducing five visual world-model agents and evaluating their frozen models. Closed‑loop rollouts reveal visual and dynamical failures such as ball disappearance and incorrect motion, while zero‑shot model‑based RL policies trained entirely within the frozen models perform poorly compared to the original agents. To address these issues, the authors introduce Concept‑Guided Spatial Regularization (CGSReg), an auxiliary loss focused on task‑critical ball regions, which improves both pixel‑space zero‑shot MBRL performance and closed‑loop rollouts for several agents.
By Yukuan Lu, Zaishuo Xia, Weyl Lu, Yubei Chen
arXiv:2606. 28128v1 Announce Type: cross Abstract: Video generation models have emerged as a promising paradigm for embodied world simulation.
By Peiwen Zhang, Yufan Deng, Shangkun Sun, Juncheng Ma, Duomin Wang, Jonas Du, Zilin Pan, Ye Huang, Hao Liang, Songyan Huang, Ruihua Zhang, Enze Xie, Ming-Yu Liu, Daquan Zhou
arXiv:2609.37250v1 Announce Type: cross
Abstract: World-action models (WAMs) couple future visual-state prediction with action generation. By adapting video generators or image-editing models pretrai...
By Yang Zhang, Jiangyuan Zhao, Chenyou Fan, Jiayu Hu, Xiu Yuan, Chenjia Bai, Xiu Li
arXiv:2609.37907v1 Announce Type: new
Abstract: Video games offer scalable environments for studying perception and control in embodied agents.Abundant online gameplay videos could supply demonstrati...
By Abhishek Pillai, Ekta Prashnani, Joohwan Kim, Iuri Frosio
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2602. 13977v2 Announce Type: replace-cross Abstract: Reinforcement learning (RL) promises to unlock capabilities beyond imitation learning for Vision--Language--Action (VLA) models, but its requirement for massive real-world interaction prevents direct deployment on physical robots.
By Zhennan Jiang, Shangqing Zhou, Yutong Jiang, Zefang Huang, Mingjie Wei, Yuhui Chen, Tianxing Zhou, Zhen Guo, Hao Lin, Quanlu Zhang, Yu Wang, Haoran Li, Chao Yu, Dongbin Zhao