For the first time, developers can also instruct the text-to-speech model to speak in a specific way—for example, “talk like a sympathetic customer service agent”—unlocking a new level of customization for voice agents.
The paper introduces a validation‑gated audit framework for voice AI customer‑care systems, treating them as stateful, multi‑turn, tool‑mediated interactions where bias and safety can manifest as added burdens before a final decision. The framework distinguishes between native speech‑to‑speech, cascaded ASR‑to‑LM‑to‑TTS, and hybrid architectures, and applies matched service facts across controlled caller presentation conditions to validate fact invariance, presentation cues, artifacts, and acoustic measurements. It outlines seven validation gates, a six‑family metric set, and demonstrates the approach with a synthetic refund‑dispute audit example, while noting that production results are withheld until the protocol is satisfied.
By Vignesh Ethiraj, Ashwath David
arXiv:2607. 21180v1 Announce Type: new Abstract: Recent advances have introduced speech-to-speech (S2S) conversational assistants capable of producing natural-sounding interactions, including non-verbal cues like tonality and mood.
By Gregor Endler, Sebastian Kraus, Lukas Stappen
arXiv:2608.28518v1 Announce Type: cross
Abstract: We investigate whether automatic speech recognition (ASR) errors in user input can lead to unsafe outputs from Embodied AI (EAI) models. We find that...
By Sihan Jia, Oliver Lemon
We’re sharing lessons from a small scale preview of Voice Engine, a model for creating custom voices.
Voxtral TTS: A frontier, open-weights text-to-speech model that’s fast, instantly adaptable, and produces lifelike speech for voice agents.
arXiv:2609.35952v1 Announce Type: cross
Abstract: We introduce HEAR (Human-recorded Evaluation of Audio-LLM bias by Real speakers), a large-scale, ecologically valid benchmark comprising 87k real hum...
By Shen Yan, Duc Le, Irina-Elena Veliche
We describe our latest thinking in the hope of helping other AI developers address safety and misuse of deployed models.
The paper titled "Safety Alignment Illusion: The Cross-Lingual Safety Gap in LLMs" highlights that current safety alignment training for large language models is predominantly English-centric, leading to failures in non‑English languages. It introduces INCLUDE, a multilingual benchmark with 2,604 prompts in six languages (English, Hindi, Bengali, Marathi, Tamil, and Hinglish) to measure Indian‑centric socio‑cultural biases. Evaluation of ten open‑ and closed‑source LLMs shows that Bengali models exhibit the highest bias scores among open‑source models, while English shows the lowest bias in open‑source but the highest in closed‑source models.
By Namya Bhatnagar
arXiv:2606. 03812v1 Announce Type: new Abstract: Operational safety in high-stakes domains such as industrial process control, autonomous, and safety-critical systems, demand reliable hazard identification.
By Sanjay Das, Ran Elgedawy, Ethan Seefried, Ryan Burchfield, Tirthankar Ghosal
Full-duplex speech models listen and speak at once, promising always-on assistants. Yet they must also decide when they should speak. Human listeners speak when addressed or when the speaker stops, bu...
Full‑duplex speech models can listen and speak simultaneously, but they struggle to decide when to speak. Experiments with five model families show that being addressed or encountering silence are reliable triggers, whereas cues like false facts or hazards are not. Even when models answer questions, they rarely challenge false claims or warn about danger, revealing a gap in content understanding and intervention decisions.
By Linkai Peng, Baorian Nuchged, Kaiqi Fu, Yuyang Yao