arXiv Computation and Language

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

The paper introduces a method called randomized intermediate guidance for training tandem speech-to-speech models, where a large language model (LLM) acts as a backend providing candidate responses while the user is speaking. Instead of simulating the backend’s guidance, the approach derives guidance directly from the conversation corpus, using target responses for informative guidance and randomly sampled responses to simulate irrelevant updates. Experiments on synthetic dialogues and 3.8k hours of real conversations show that this technique yields response quality comparable to LLM-generated baselines while improving natural turn‑taking and audio‑judge naturalness.

Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv Computation and Language
Sep 22

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...

By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv Computation and Language
Sep 4

Decoupling Turn-Taking from Semantics: A Decoupled Data Approach for Finite-State-Machine-Based Full-Duplex Dialogue

The paper introduces a decoupled data approach for the Neural Finite State Machine (NFSM) framework to improve full‑duplex dialogue. It serializes real human‑human spoken dialogues into FSM tapes using a rule‑based event‑guided transformation, while shaping semantics through human‑agent text dialogues. A Source‑Aware Calibrated (SAC) loss is proposed to balance state‑transition token distribution and align each data source with its strongest supervisory signal, leading to better turn‑taking performance without sacrificing semantic quality.

By Yihang Li, Chenhui Chu
arXiv Computation and Language
Sep 14

SteerDuplex: Steerable Duplex Speech Dialogue Models

SteerDuplex is a full‑duplex speech dialogue model that can be steered along attributes such as tone, persona, speaking rate, and voice style in response to user instructions. The authors introduce a taxonomy of text‑ and audio‑based steerability, identify gaps in existing models, and fine‑tune a Moshi‑based model with reinforcement learning to improve timing and response continuity. They also present SteerBench, a benchmark of 390 spoken prompts and 1,067 human‑authored rubrics, showing significant gains in audio‑steering pass rates and interruption handling compared to open baselines.

By Utkarsh Tyagi, Ramaneswaran Selvakumar, Advait Gosai, Sonal Kumar, Nikhil Barhate, Isabell Sagar, Steven Li, Miheer Bavare, Daniel Quigley, Fabiola Tapia Carrillo, Jose M Patron E, Diego Mac\'ias Guti\'errez, Paul Song, Ramani Duraiswami, Dinesh Manocha, Yunzhong He
arXiv AI
Sep 24

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.

By Marcin Sowa\'nski, Kacper Leszczy\'nski, Kacper Krzywicki, Krzysztof Wodnicki
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
arXiv AI
3d ago

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

The paper introduces Human-1, the first open and reproducible full‑duplex spoken dialogue system for Hindi. It adapts the Moshi architecture with a custom Hindi tokenizer and trains on 26,000 hours of real spontaneous conversations from 14,695 speakers, enabling the model to learn natural turn‑taking, interruptions, overlaps, and backchannels. A two‑stage training process—large‑scale pre‑training followed by fine‑tuning on 1,000 hours—yields a system that, according to both automatic metrics and human judgments, generates natural and meaningful full‑duplex conversational behavior in Hindi.

By Bhaskar Singh, Manas Dhir, Shobhit Banga, Pranav Sharma