arXiv:2606. 27326v1 Announce Type: new Abstract: Modern generative world models render increasingly realistic action-controllable futures, yet they frequently hallucinate: rollouts remain visually fluent while drifting from the ground-truth dynamics.
By Nicklas Hansen, Xiaolong Wang
WorldAgen is a unified framework that jointly learns world modeling and action prediction using a shared Transformer backbone with two specialized heads. It introduces a Mixed Unidirectional Attention Mask to separate the world model and agent model, and enables Test-Time Training (TTT) by sampling exploratory actions and updating the world model with real state transitions. Experiments on CALVIN and LIBERO show that WorldAgen matches or surpasses state‑of‑the‑art methods, especially when TTT is applied to a few samples.
By Chi Wan, Kangrui Wang, Yuan Si, Pingyue Zhang, Manling Li
arXiv:2609.39182v1 Announce Type: cross
Abstract: World Models are appearing as the next major frontier in computer vision. However, their robustness is currently largely unexplored. We identify the...
By Ali J Alrasheed, Aryan Yazdan Parast, Basim Azam, James Bailey, Naveed Akhtar
World-Coherent Decoding (WCD) is a test-time planning framework for World Action Models (WAMs) that treats rollouts as falsifiable future–action hypotheses. At each decision step, WCD samples multiple candidates from a frozen WAM and ranks them using flow-based video surprisal for visual plausibility and action path effort for generation stability. After execution, the observed outcome audits the chosen imagination, producing a mismatch signal that trains a lightweight online predictor to improve future candidate selection, thereby enhancing reliability without updating the backbone model.
By Chuhan Zhang, Seiji Ito, Kenta Hoshino, Satoshi Ikehata, Ikuro Sato
arXiv:2607. 14180v1 Announce Type: cross Abstract: World models are widely used in offline reinforcement learning (RL) to improve sample efficiency and generate experience beyond a fixed dataset.
By Logan Mondal Bhamidipaty, Mykel Kochenderfer, Subramanian Ramamoorthy
arXiv:2609.39235v1 Announce Type: cross
Abstract: World models offer a promising way to help robots understand how the physical world evolves and plan complex behaviours through imagination. Yet exis...
By Ali Alrasheed, Basim Azam, Naveed Akhtar
arXiv:2608. 11605v1 Announce Type: new Abstract: World Action Models (WAMs) couple future visual prediction with robot action generation, enabling policies to model how the physical world evolves during interaction.
By Jiakai Huang, Zhongbo Wu, Zheng Zhang, Zihan Wang, Shan You, Tao Huang
arXiv:2606. 07974v1 Announce Type: cross Abstract: A learned world model provides a powerful physical intuition for evaluating future states.
By Yuhai Wang, Jiawei Xia, Rongxuan Zhou, Xiao Hu, Yongliang Shi, Jing Du, Yang Ye
arXiv:2606. 26217v1 Announce Type: new Abstract: Joint-Embedding Predictive Architectures (JEPAs), including recent LeWorldModel (LeWM), have become a promising foundation for reconstruction-free visual world models.
By Yuntian Gao, Xiangyu Xu
arXiv:2607. 07993v1 Announce Type: cross Abstract: Identifying faithfulness hallucinations in LLM-generated outputs remains challenging due to the scarcity of high-quality annotated data.
By Shiping Yang, Shining Liang, Weihao Liu, Wenbiao Ding, Linjun Shou, Lu Cheng, Angel X. Chang
arXiv:2602. 01740v3 Announce Type: replace Abstract: Video language models (Video-LLMs) are prone to hallucinations, generating plausible but ungrounded content when visual evidence is weak, ambiguous, or biased.
By Qixin Xiao, Kun Zhou
arXiv:2609.34981v2 Announce Type: replace
Abstract: World action models (WAMs) predict the future alongside actions during training. Due to the heavy computation cost of video denoising, whether the...
By Renping Zhou, Zanlin Ni, Zihao Fan, Guohao Fu, Zeyu Liu, Hao Shi, Jie Zhang, Chi Bene Chen, Yang Yue, Xueyang Fu, Gao Huang