DeepMind Blog

Advanced audio dialog and generation with Gemini 2.5

Read the original on DeepMind Blog →

Gemini 2. 5 has new capabilities in AI-powered audio dialog and generation.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at DeepMind Blog.

Simon Willison
Sep 23

Gemini 3.8 TTS Playground

Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.