Sora 2 is here
Our latest video generation model is more physically accurate, realistic, and controllable than prior systems. It also features synchronized dialogue and sound effects.
Sora 2 is our new state of the art video and audio generation model. Building on the foundation of Sora, this new model introduces capabilities that have been difficult for prior video models to achieve– such as more accurate physics, sharper realism, synchronized audio, enhanced steerability, and an expanded stylistic range.
Our latest video generation model is more physically accurate, realistic, and controllable than prior systems. It also features synchronized dialogue and sound effects.
Sora is OpenAI’s video generation model, designed to take text, image, and video inputs and generate a new video as an output. Sora builds on learnings from DALL-E and GPT models, and is designed to give people expanded tools for storytelling and creative expression.
Our video generation model, Sora, is now available to use at sora. com.
To address the novel safety challenges posed by a state-of-the-art video model as well as a new social creation platform, we’ve built Sora 2 and the Sora app with safety at the foundation. Our approach is anchored in concrete protections.
To address the novel safety challenges posed by a state-of-the-art video model as well as a new social creation platform, we’ve built Sora 2 and the Sora app with safety at the foundation. Our approach is anchored in concrete protections.
Filmmaker Lyndon Barrois describes how to use Sora as a storytelling tool.
The paper "Less is More: Encoder-only Audio-Visual Segmentation" introduces EASE, an encoder-only model for Audio‑Visual Semantic Segmentation (AVSS). EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—about three times faster than previous Transformer‑based AVSS models—and trains in under 11 GPU‑hours. The authors demonstrate that simpler, faster architectures can match or exceed the performance of more complex models across various backbones and resolutions.
The paper introduces EASE, an encoder‑only model for Audio‑Visual Semantic Segmentation that eliminates redundant components found in prior Transformer‑based approaches. EASE achieves state‑of‑the‑art accuracy while running at up to 365 FPS—three times faster than previous models—and can be trained in under 11 GPU‑hours. The authors provide code, weights, and samples, positioning EASE as a scalable foundation for future research and real‑time applications.
arXiv:2609.13830v1 Announce Type: new Abstract: We present DiVA, a deeply interactive digital life simulator pioneering a new paradigm for long-term, open-ended interactive experiences within digital...
arXiv:2609.31810v2 Announce Type: replace-cross Abstract: Video-to-music models have advanced considerably in the last few years, particularly in film and music video applications. In this paper, we...
arXiv:2607. 03118v1 Announce Type: cross Abstract: We introduce Vidu S1, a real-time interactive video generation model supporting voice control of digital characters.
Since we introduced Sora to the world last month, we’ve been working with artists to learn how Sora might aid in their creative process.