Advancing voice intelligence with new models in the API
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.
Explore new realtime voice models in the OpenAI API that can reason, translate, and transcribe speech, enabling more natural and intelligent voice experiences.
Developers can now build fast speech-to-speech experiences into their applications
Exploring the technology behind our text-to-speech model.
We’re sharing lessons from a small scale preview of Voice Engine, a model for creating custom voices.
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.
Parloa leverages OpenAI models to power scalable, voice-driven AI customer service agents, enabling enterprises to design, simulate, and deploy reliable, real-time interactions.
A new generation of voice models for natural human-AI interaction, now powering ChatGPT Voice.
GPT‑Live‑1 brings natural, full-duplex voice conversations to the API, with stronger instruction following, custom voices, and telephony support.
Conversation Coach is a voice‑first AI system that lets managers rehearse difficult workplace conversations in a realistic spoken format. It tackles low‑latency interaction, adaptive bot personalities that simulate various employee types, and personalized feedback on content and policy compliance. The authors compare an end‑to‑end speech‑to‑speech model with a cascaded approach, finding the former offers lower latency and cost, while the cascaded model provides better reasoning for coaching quality, and they deployed the cascaded architecture to 40,000+ managers over six months.
arXiv:2607. 21180v1 Announce Type: new Abstract: Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood.
Retell AI is transforming the call center with AI voice automation powered by GPT-4o and GPT-4. 1.