arXiv:2608.30692v1 Announce Type: new
Abstract: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world. W...
By Joonghyuk Shin, Yicong Hong, Jaesik Park, Xun Huang
HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.
By Haoyun Tang, Haodong Cui, Keyao Xu, Zhandong Mei, Kun Wang
arXiv:2609.39324v1 Announce Type: cross
Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. Howeve...
By Jingqiu Wang, Yan Wang
arXiv:2608.29904v1 Announce Type: new
Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
By Hai Nguyen-Truong, Tuan-Anh Vu, Dang Huynh
arXiv:2608. 12078v1 Announce Type: cross Abstract: Learning world models from offline trajectories enables agents to accomplish different tasks through planning.
By Shukrullo Nazirjonov, Sai Prasanna, Anna Manasyan, Georg Martius
World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.
By Chenxi Song, Yanming Yang, Chi Zhang
arXiv:2609.13006v2 Announce Type: replace
Abstract: Video diffusion models (VDMs) synthesize photorealistic content, yet they often fail to follow the course that a physical phenomenon should take wi...
By Minh-Loi Nguyen, Xuan-Vu Le, Trung-Nghia Le, Tam V. Nguyen, Minh-Triet Tran, Thanh-Toan Do
The paper introduces Recency Forcing, a technique that addresses the long‑horizon degradation in autoregressive video generation caused by KV eviction mismatch. By applying a timestep‑dependent bias—Temporal Response Bias—derived from a positional response measure, the method reduces the influence of distant frames during inference without altering context length or training objectives. An exact reformulation, Biased Attention Reparameterization, enables this bias to be applied as a standard FlashAttention call with zero overhead, achieving state‑of‑the‑art long‑horizon generation quality on VBench datasets.
By Tri Cao, Hung Nguyen, Phong Nguyen, Khoi Nguyen
arXiv:2608. 07408v1 Announce Type: cross Abstract: We study visual persistence in interactive video world models.
By Xindi Wu, Sven Elflein, James Lucas, Olga Russakovsky, Laura Leal-Taix\'e, Despoina Paschalidou, Jonathan Lorraine, Aljo\v{s}a O\v{s}ep
The paper introduces Action-Contrastive Masked Transition Modeling (AC‑MTM), a method that stabilizes Joint‑Embedding Predictive Architectures (JEPAs) without relying on Gaussian regularization. AC‑MTM adds a training‑only inverse‑dynamics head that uses Action‑NCE to force each latent transition to identify its generating action, thereby preventing encoder collapse. Experiments on pixel‑control and multi‑object visual tasks show that AC‑MTM trains stably from scratch and matches or surpasses the performance of SIGReg, achieving up to a 24‑point improvement on the OGBench Visual Scene benchmark.
arXiv:2609.39563v1 Announce Type: new
Abstract: Existing video language models encode sampled RGB frames independently, so a long video must either exhaust the token budget or drop the changes betwee...
By Can Zhang, Xiaotian Han, Junyuan Shang, Yuchen Ding, Zhenyu Zhang, Shuohuan Wang, Dianhai Yu, Ruirui Li
arXiv:2608.23526v1 Announce Type: new
Abstract: World models can predict video without learning dynamics that they reliably preserve. We test whether a frozen DreamerV3 trained only on pendulum video...
By Richard Bao