Simon Willison

Gemini Live audio

Google has launched Gemini 3.8 Live and Gemini 3.8 Live Extended Thinking, two speech‑to‑speech models similar to OpenAI’s GPT‑Live family. A web UI built with GPT‑6 Astra Extra High lets users select a model, choose a voice preset, provide an optional system prompt, and engage in voice conversations directly in the browser, even interrupting the model while it speaks. The implementation relies on a WebSocket endpoint and the Web Audio API for capturing and playing audio, with no external libraries required.

Simon Willison
Sep 23

Gemini 3.8 TTS Playground

Google has launched two new Gemini text‑to‑speech models—gemini‑3.8‑flash‑tts and gemini‑3.8‑flash‑lite‑tts—offering a library of over 2,000 voices and the option to create a custom voice from a 30‑second audio sample. The author built a playground interface that lets users define multi‑character conversations with distinct voices and styles, and demonstrated it with a scripted dialogue between two pelicans. Generating 1 minute 18 seconds of audio with the Flash model took about 20 seconds and cost 2.74 cents.

arXiv AI
Sep 21

NemotronLabs VoiceChat: An Open Full-duplex Speech-to-Speech Model with Tool Calling Capabilities

arXiv:2609.21967v1 Announce Type: cross Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combine...

By Jagadeesh Balam, Travis Bartley, Edresson Casanova, Sanjay Chauhan, Chen Chen, Zhehuai Chen, Zijia Chen, Francesco Ciannella, Slyne Deng, Mikyas Desta, Harishchandra Dubey, Slim Essid, Nourchene Ferchichi, Boris Ginsburg, Mariana Graterol Fuenmayor, Negar Habibi, Kevin Hu, Anand Joseph, Viraj Karandikar, Myungjong Kim, Viacheslav Klimkov, Seelan Lakshmi Narasimhan, Lily Lee, Jason Li, Eileen Long, Ameya Mahabaleshwarkar, Aditya Malte, Adi Margolin, Sasha Meister, Valentin Mendelev, Oluwatobi Olabiyi, Ankita Pasad, Yifan Peng, Elena Rastorgueva, Jayda Ritchie, Jason Roche, Nikhil Srihari, Yuanhang Su, Yoshi Suhara, Viet Anh Trinh, Jinhan Wang, Piotr Zelasko, Hui Wang, Puhui Meng, Chaosen Zhang, Yunsheng Liu, Shawn Wang, Wenjing Li, Zhonglei He
arXiv Computer Vision
Sep 15

Realtime-Venus: A full-duplex interaction system with asynchronous delegation

arXiv:2609.13814v1 Announce Type: new Abstract: Natural interaction in digital and physical environments requires continuous perception and timely responses. Spoken dialogue relies on acoustic and li...

By Ruixiang Zhao, Hualei Wang, Renhe Sun, Enzhi Zhou, Jincenzi Wu, Xujie Song, Kexin Shi, Zihang Liu, Pengcheng Zhu, Jiayi Zhou, Baoyue Zhang, Changhao Zhang, Zitong Wang, Jinhong Wang, Tong Niu, Jingjing Liu, Junan Lin, Haolin He, Hengshuo Chu, Yuhui Chen, Jian Liu, Yuge Huang, Junliang Xing, Yuntao Wang, Weiqiang Wang, Chun Yu, Yuanchun Shi
arXiv Computation and Language
Sep 21

I'll Keep an Ear Out: Teaching AudioLLMs Proactive Audio Assistance

arXiv:2609.21183v1 Announce Type: cross Abstract: Audio large language models (AudioLLMs) operate reactively, responding only when queried. We introduce proactive audio assistance, where an AudioLLM...

By Amit Kumar Singh Yadav, Ritvik Shrivastava, Xuan Zhang, Seungwhan Moon, Shashank Jain, Pinar Donmez, Babak Damavandi
arXiv AI
Aug 28

From Sound to Symptom: Real-Time Respiratory Signal Understanding for Conversational Healthcare Agents

The paper introduces HealthCUES, a real‑time streaming pipeline that extracts and analyzes cough and throat‑clearing events from live spoken conversations. It detects coughs within sub‑second latency, distinguishes cough subtypes (dry, wet, barking, whooping), differentiates coughing from throat clearing, and estimates temporal boundaries, all while gating alerts based on conversational context. The system, built on Qwen3Omni, achieves high accuracy (93% F1 for cough detection) and low latency (340 ms) and has been validated by healthcare professionals for telehealth use.

By Tanmay Laud, Herprit Mahal, Subhabrata Mukherjee