Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
Tolan built a voice-first AI companion with GPT-5. 1, combining low-latency responses, real-time context reconstruction, and memory-driven personalities for natural conversations.
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4. 1.
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
We’re upgrading the GPT-5 series with warmer, more capable models and new ways to customize ChatGPT’s tone and style. GPT-5.
GPT-5. 2 is our most advanced frontier model for everyday professional work, with state-of-the-art reasoning, long-context understanding, coding, and vision.
Invideo AI uses OpenAI’s GPT-4. 1, gpt-image-1, and text-to-speech models to transform creative ideas into professional videos in minutes.
arXiv:2509. 00078v2 Announce Type: replace-cross Abstract: The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question.