arXiv:2608. 04964v1 Announce Type: new Abstract: Interactive video world models are essential for long-horizon planning and exploration, yet they suffer from compounding errors.
By Bohai Gu, Yueyang Yuan, Taiyi Wu, Dazhao Du, Jian Liu, Xiaoyi Pang, Jie Zhang, Xiaocheng Lu, Haobin Zhong, Xiaotong Zhao, Alan Zhao, Song Guo
arXiv:2608. 10232v1 Announce Type: cross Abstract: Recent world-action models (WAMs) show that co-training policies with future prediction can provide physical priors for action generation.
By Quanquan Peng, Yutong Liang, Rui Yan, Nicklas Hansen, Xiaolong Wang
arXiv:2607. 28362v1 Announce Type: cross Abstract: We present ShadowDancer, a novel approach to any-action, frame-level control of interactive video world models.
By Jin Cao, Zian Meng, Kaipeng Zhang
arXiv:2606. 24152v1 Announce Type: cross Abstract: Existing literature claims that video generation essentially is world modelling.
By Xin Wang, Wenxuan Liu, Tongtong Feng, Wenwu Zhu
arXiv:2608. 15156v1 Announce Type: cross Abstract: World models may predict the future without making clear which parts of their hidden state actually drive those predictions.
By Yang Liu, Yuming Chen
Autoregressive video generation has emerged as a powerful paradigm for World Action Models (WAMs). However, existing approaches suffer from slow training convergence and limited converged accuracy, particularly at high frame rates, as the training supervision is confined to the current chunk without explicit signals about future dynamics; they also suffer from slow inference due to iterative video denoising.