DeepMind Blog

Gemini 3.1 Flash Live: Making audio AI more natural and reliable

Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.

arXiv Machine Learning
Jul 21

ChipChat: Low-Latency Cascaded Conversational Agent in MLX

arXiv:2509. 00078v2 Announce Type: replace-cross Abstract: The emergence of large language models (LLMs) has transformed spoken dialog systems, yet the optimal architecture for real-time on-device voice agents remains an open question.

By Tatiana Likhomanenko, Richard He Bai, Zijin Gu, Zakaria Aldeneh, Shiladitya Dutta, Luke Carlson, Han Tran, Yizhe Zhang, Ruixiang Zhang, Huangjie Zheng, Navdeep Jaitly
arXiv AI
Jul 22

What the Waveform Knows: Transparent-first Speech and Audio Intelligence with Caption Studio

arXiv:2607. 18704v1 Announce Type: cross Abstract: Caption Studio is a transparency-first speech and audio intelligence platform that transforms spoken audio and video into structured, searchable content through automated transcription, speaker diarization, speech analytics, signal-level audio analysis, and subtitle generation.

By Cheng Siong Chin, Jianhua Zhang, Mohan Venkateshkumar