Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
Tolan built a voice-first AI companion with GPT-5. 1, combining low-latency responses, real-time context reconstruction, and memory-driven personalities for natural conversations.
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4. 1.
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
Conversation Coach is a voice‑first AI system that lets managers rehearse difficult workplace conversations in a realistic spoken format. It tackles low‑latency interaction, adaptive bot personalities that simulate various employee types, and personalized feedback on content and policy compliance. The authors compare an end‑to‑end speech‑to‑speech model with a cascaded approach, finding the former offers lower latency and cost, while the cascaded model provides better reasoning for coaching quality, and they deployed the cascaded architecture to 40,000+ managers over six months.
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.
We’re upgrading the GPT-5 series with warmer, more capable models and new ways to customize ChatGPT’s tone and style. GPT-5.