Oneira: From Open-Ended Generation to Open-World Interaction in Video World Models
Read the original on arXiv Computer Vision →The Flow has not summarised this story yet — read it at arXiv Computer Vision.
The Flow has not summarised this story yet — read it at arXiv Computer Vision.
arXiv:2610.02162v1 Announce Type: new Abstract: How can a world model continuously observe regions beyond the actor's current view? Video world models simulate how an environment evolves from an agen...
The paper introduces CoDeR, a new paradigm for world modeling that explicitly builds an executable world using code rather than relying solely on visual observations. CoDeR translates high‑level concepts into structured world rules, executable dynamics, and perceptual observations through five complementary roles, enabling long‑term memory, open‑ended interactions, autonomous world evolution, and persistent multi‑agent dynamics. Experiments show that this framework extends the capabilities of existing world models and achieves state‑of‑the‑art performance across multiple evaluation settings.
World in World introduces a training‑free, inference‑time interface that transforms diverse control signals—such as source‑video observations, target‑view projections, geometry renderings, and retrieved states—into camera‑ and time‑labelled visual states. These states are processed by a frozen causal video model’s self‑attention, enabling tasks like camera‑controlled rerendering, long‑horizon revisiting, and human‑motion transfer without additional training. The method employs a correspondence router and evidence‑wise attention to align token identities and regulate auxiliary channel contributions during a single denoising pass.
GameWAM is the first World-Action Model designed for native closed-loop gameplay and GUI control in modern video games. It jointly generates future visual observations and executable keyboard-mouse trajectories using parallel visual and action generative processes, block-causal conditioning, and flow matching. The model predicts gameplay/GUI mode at each step, handles heterogeneous native controls, and employs block-cycle control for long-horizon interaction, achieving competitive task success with fewer native actions than prior agents.
arXiv:2606. 02753v1 Announce Type: cross Abstract: Video world models are a foundational generative technology for embodied AI and the Metaverse, yet existing approaches are inherently limited to a single agent observing from a single perspective.
The paper introduces Code World Model, a framework that decouples world evolution from visual rendering by using a coding agent as a world brain. The agent reasons about events, generates executable code to maintain persistent state, and a proxy representation links this state to a video model for high‑fidelity visual output. Experiments with MiniMax‑H3 show that the system can follow proxy‑based spatiotemporal specifications while preserving rich visual dynamics, illustrating a new approach to open‑ended world modeling.