Introducing GPT-Live
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
GPT-Live enables continuous voice interaction with AI, using a turnless speech model and low-latency architecture for faster, more natural conversations.
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4. 1.
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
Developers can now build fast speech-to-speech experiences into their applications
Tolan built a voice-first AI companion with GPT-5. 1, combining low-latency responses, real-time context reconstruction, and memory-driven personalities for natural conversations.
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
How OpenAI rebuilt its WebRTC stack to power real-time Voice AI with low latency, global scale, and seamless conversational turn-taking.
arXiv:2509. 00078v2 Announce Type: replace-cross Abstract: The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question.
The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.
Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and linguistic cues, while video interaction also requi...