← Back to all news
OpenAI Blog September 21, 2022

Introducing Whisper

Read the original on OpenAI Blog →

The Flow has not summarised this story yet — read it at OpenAI Blog.

  • multimodal

Related stories

OpenAI Blog
Apr 24, 2024

Introducing ChatGPT and Whisper APIs

llmsmultimodal
More like this →
arXiv AI
Aug 12

Whisper-Aware LLM: Self-Supervised Uncertainty Learning for Robust Whispered Speech Recognition

arXiv:2608. 10836v1 Announce Type: cross Abstract: The signal ambiguity of whispered speech drives ASR systems toward two opposing failure modes: failing to capture whispered speech or hallucinatory transcription of noise.

By Gaopeng Xu, Zhenyu Wang, Zheng Xue, Yinfeng Xia, Haitao Yao
llmsmultimodalbenchmarkssafety
More like this →
OpenAI Blog
Mar 20, 2025

Introducing next-generation audio models in the API

For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.

agentsmultimodal
More like this →
OpenAI Blog
Sep 25, 2023

ChatGPT can now see, hear, and speak

llms
More like this →
DeepMind Blog
Apr 15

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

llms
More like this →
OpenAI Blog
Oct 1, 2024

Introducing the Realtime API

Developers can now build fast speech-to-speech experiences into their applications

More like this →