MOSH-WM: Mask-Grounded Soft-Hamiltonian Dynamics for Object-Centric World Models
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2608.30692v1 Announce Type: new Abstract: Video world models are increasingly used as simulators, yet visual fidelity alone does not show that a model maintains the hidden state of the world. W...
HaM-World introduces a structured world model that combines history-conditioned selective memory with a Soft‑Hamiltonian latent dynamics prior. The model decomposes the latent state into a canonical (q,p) subspace governed by an energy‑derived Hamiltonian vector field and a context subspace c capturing non‑conservative factors, while Mamba selective state‑space memory conditions the transition used for prediction, reward, value estimation, and planning. Across six DeepMind Control Suite tasks, HaM-World achieves top rankings on four tasks, improves average AUC, reduces imagined‑rollout error by 45% on short‑to‑medium horizons, and outperforms baselines under 12 out‑of‑distribution perturbations.
arXiv:2609.39324v1 Announce Type: cross Abstract: Vision-Language-Action (VLA) models have recently incorporated world models to provide richer dynamic supervision beyond sparse action labels. Howeve...
arXiv:2608.29904v1 Announce Type: new Abstract: Modern video generators routinely fail at physical dynamics: objects float, trajectories violate gravity, contacts vanish. Standard denoising and flow-...
arXiv:2608. 12078v1 Announce Type: cross Abstract: Learning world models from offline trajectories enables agents to accomplish different tasks through planning.
World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.