OpenAI Blog

Introducing gpt-realtime and Realtime API updates

We’re releasing a more advanced speech-to-speech model and new API capabilities including MCP server support, image input, and SIP phone calling support.

arXiv AI
Jul 22

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

arXiv:2607. 18704v1 Announce Type: cross Abstract: Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searchable content through automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle generation.

By Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar
Hugging Face Trending Papers
Jul 20

Re-Sonance: A Dysarthric Asynchronous Real-Time Speech Conversion System Based on a Three-Stage Cascaded ASR-LLM-TTS Architecture

Individuals with dysarthria face significant challenges in professional speaking scenarios such as conferences, presentations, and meetings, where real-time communication is crucial. While existing Augmentative and Alternative Communication (AAC) systems provide basic support, they often fail to meet the demands of professional speaking environments due to high latency and unnatural speech patterns.