DeepMind Blog

Gemini 3.1 Flash TTS: the next generation of expressive AI speech

Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.

Simon Willison
Sep 15

Gemini Live audio

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech‑to‑speech models similar to OpenAI’s GPT‑Live family. A web UI built with GPT‑6 Astra Extra High lets users select a model, choose a voice preset, provide an optional system prompt, and engage in voice conversations directly in the browser, even interrupting the model while it speaks. The implementation relies on a WebSocket endpoint and the Web Audio API for capturing and playing audio, with no external libraries required.

Simon Willison
Sep 23

Gemini 3.8 TTS Playground

Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.