Project Genie: Experimenting with infinite, interactive worlds
Google AI Ultra subscribers in the U. S.
Genie 3 can generate dynamic worlds that you can navigate in real time at 24 frames per second, retaining consistency for a few minutes at a resolution of 720p.
Google AI Ultra subscribers in the U. S.
arXiv:2607. 18367v1 Announce Type: new Abstract: Unlike conventional video game development, which relies on labor-intensive pipelines for asset production, animation, physics, and programming, video world models generate interactive environments from user inputs instantly.
arXiv:2606. 30292v1 Announce Type: new Abstract: We present DreamForge-World 0.
arXiv:2603. 03482v2 Announce Type: replace-cross Abstract: Interactive world models continually generate video by responding to a user's actions, enabling open-ended generation capabilities.
We introduce PhiZero, a physical world model built around physical language, a compact discrete representation of world-state transitions. Existing physical world models typically predict future videos directly in pixel space, leaving the underlying world dynamics implicit within high-dimensional visual predictors.
HelloWorld is a 2B-parameter driving world model that learns from diverse video data to generate controllable driving scenarios. It uses ego pose, HD maps, and 3D boxes to produce coherent multi‑sensor outputs, including synchronized seven‑camera RGB and conditional LiDAR. The system is designed for efficient repeated inference and is evaluated on visual quality, control fidelity, cross‑view consistency, robustness, and inference speed.
arXiv:2508.13009v5 Announce Type: replace Abstract: Recent advances in interactive video generations have demonstrated diffusion model's potential as world models by capturing complex physical dynami...
arXiv:2607. 06401v1 Announce Type: new Abstract: World models -- internal simulators that learn the structure and dynamics of an environment -- have become one of the most actively debated concepts in AI.
arXiv:2607.26694v3 Announce Type: replace Abstract: We present Visko Orbis 1.0, a Live Model for real-time, interactive long video generation. Users can change the prompt at any moment during generat...
arXiv:2603. 02697v2 Announce Type: replace-cross Abstract: This paper presents ShareVerse, a video generation framework enabling multi-agent shared world modeling, addressing the gap in existing works that lack support for unified shared world construction with multi-agent interaction.
arXiv:2610.01614v1 Announce Type: new Abstract: Generative video world models can now synthesize open-ended environments that agents can navigate and interact with in simple ways. Yet open-ended gene...
GEM is a Generative LiDAR world model that uses a deformable Mamba architecture to better handle the disorder of LiDAR point clouds and distinguish dynamic objects from static structures. The model tokenizes LiDAR sweeps, unsupervisedly disentangles dynamic and static features, and applies a tri‑path deformable Mamba for selective scanning and adaptive gating fusion, improving spatial‑temporal understanding. Experiments show GEM outperforms existing methods across multiple benchmarks, and it can be paired with a planner and BEV controller for autonomous rollout and "what‑if" scenario generation.