Gemini 3.1 Flash Live: Making audio AI more natural and reliable
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
Our latest voice model has improved precision and lower latency to make voice interactions more fluid, natural and precise.
Gemini 2. 5 has new capabilities in AI-powered audio dialog and generation.
Our newest audio model introduces granular audio tags that give you precise control to direct AI speech for expressive audio generation.
Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech‑to‑speech models similar to OpenAI’s GPT‑Live family. A web UI built with GPT‑6 Astra Extra High lets users select a model, choose a voice preset, provide an optional system prompt, and engage in voice conversations directly in the browser, even interrupting the model while it speaks. The implementation relies on a WebSocket endpoint and the Web Audio API for capturing and playing audio, with no external libraries required.
DeepMind has introduced Gemini 3.5 Transcribe, a new tool that offers more intelligent speech-to-text transcription. The update promises improved accuracy and smarter handling of spoken content, enhancing the overall transcription experience.
Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.
arXiv:2606. 17126v1 Announce Type: cross Abstract: Singing style is a crucial aspect of a natural and expressive singing voice.
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.
arXiv:2607. 07985v1 Announce Type: cross Abstract: We report the empirical reliability of Gemini models as audio judges that score full-duplex agent conversations directly from the raw stereo waveform, tested across three models in the Gemini family: 2.
arXiv:2604. 14606v2 Announce Type: cross Abstract: Universal speech enhancement (USE) aims to restore speech signals from diverse distortions across multiple sampling rates.
We’re sharing lessons from a small scale preview of Voice Engine, a model for creating custom voices.
Gemini 3. 5 Live Translate brings near real-time, natural speech translation to Google AI Studio, Google Translate and Google Meet.