Introducing Whisper
Related stories
Introducing ChatGPT and Whisper APIs
Blazingly fast whisper transcriptions with Inference Endpoints
Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition
arXiv:2608. 10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise.
Chat Templates: An End to the Silent Performance Killer
Introducing next-generation audio models in the API
For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.
ChatGPT can now see, hear, and speak
Gemini 3.1 Flash TTS: the next generation of expressive AI speech
Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
Introducing the Realtime API
Developers can now build fast speech-to-speech experiences into their applications
Introducing Le Chat Enterprise
Introducing gpt-realtime and Realtime API updates
We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.