MultiTalk: Scaling Full-Duplex Speech Models to Long, Multi-Party, Bilingual Conversation
Read the original on arXiv AI →The Flow has not summarised this story yet — read it at arXiv AI.
The Flow has not summarised this story yet — read it at arXiv AI.
DuplexDrama is a newly announced synthesized spoken dialogue dataset that uniquely combines complete persona and scenario settings, three full‑duplex behaviors (interruption, backchannel, incomplete), expressive speech with persona‑aligned emotion labels, and script‑aware sound events. The dataset was created through a four‑stage pipeline and validated for quality on both scripts and audio, yielding over 2,000 hours of audio featuring 64 voices across 13 personas and 5 age groups, with 3.8% of turns containing full‑duplex behaviors. A curated bilingual subset of 6,400 dialogues (800 hours total) will be released to support research in full‑duplex spoken dialogue models, and evaluation prompts will accompany the dataset.
SteerDuplex is a full‑duplex speech dialogue model that can be steered along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. The authors introduce a taxonomy of text‑ and audio‑based steerability, identify gaps in existing models, and fine‑tune a Moshi‑based model with reinforcement learning to improve timing and response continuity. They also present SteerBench, a benchmark of 390 spoken prompts and 1,067 human‑authored rubrics, showing significant gains in audio‑steering pass rates and interruption handling compared to open baselines.
arXiv:2606.11167v2 Announce Type: replace Abstract: Full-duplex spoken dialogue models can listen and speak simultaneously, making them a promising architecture for natural conversation. However, cur...
arXiv:2609.24812v1 Announce Type: new Abstract: Voice provides a natural and immediate interface for AI agents. Many settings in which voice agents could be useful, including meetings, households, an...
arXiv:2609.21967v1 Announce Type: cross Abstract: We introduce NemotronLabs VoiceChat, an open full-duplex speech-to-speech model with native tool-calling capabilities. NemotronLabs VoiceChat combine...
Full-duplex speech models require training data that preserves turn-taking, overlap, interruption, and backchannel behavior, yet these signals are entangled across speakers in noisy real-world recordi...