arXiv AI

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.

arXiv Computation and Language
3d ago

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

The paper introduces a method called randomized intermediate guidance for training tandem speech-to-speech models, where a large language model (LLM) acts as a backend providing candidate responses while the user is speaking. Instead of simulating the backend’s guidance, the approach derives guidance directly from the conversation corpus, using target responses for informative guidance and randomly sampled responses to simulate irrelevant updates. Experiments on synthetic dialogues and 3.8k hours of real conversations show that this technique yields response quality comparable to LLM-generated baselines while improving natural turn‑taking and audio‑judge naturalness.

By Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv Machine Learning
Sep 21

Talk to Me, Jarvis: An Open-Source Edge-Deployable Voice Assistant Framework for Autonomous Racecars

The paper introduces Jarvis, an offline, edge‑deployable voice assistant designed for autonomous racecars. It combines speech recognition, synthesis, and a lightweight text‑to‑command classifier fine‑tuned from the Mistral 7B model to provide high‑level behavioral commands. Experiments show 97.63 % intent recognition accuracy with an average latency of 1.39 s, outperforming larger online‑hosted models and enabling quick response times for time‑critical driving tasks.

By Daniel Henel, Frederik Werner, Alexander Langmann, Johannes Betz
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
Hugging Face Trending Papers
Jun 11

Adaptive Turn-Taking for Real-time Multi-Party Voice Agents

Turn-taking in multi-party spoken conversations remains a fundamental challenge for voice-based agents, particularly under dynamic floor competition and varying user expectations. We propose ModeratorLM, a role-playing voice agent that conditions turn-taking behavior on an explicitly assigned role in multi-party settings.

arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv Computation and Language
Sep 22

COT-TTS: Audio Context-Aware Text-to-Speech with Chain-of-Thought Reasoning

arXiv:2609.22697v1 Announce Type: new Abstract: Recently, text-to-speech systems have made significant progress in speech expressiveness and controllability. However, the speaking style of generated...

By Weizhen Bian, Sitong Cheng, Rongxiu Zhong, Jiahao Pan, Liumeng Xue, Boyi Kang, Shilei Zhang, Jinglei Liu, Yue Wang, Junlan Feng, Bei Liu, Wei Xue
arXiv Computation and Language
Sep 14

Not All Speech Is Intent: Adaptive Self-Correcting Inference Layer for Post-ASR False Wake-Up

The paper introduces ASCIL, a post‑ASR correction framework that re‑evaluates wake‑up intent by combining acoustic embeddings, linguistic cues, device context, and past misclassifications. ASCIL interprets both implicit (hesitation, disengagement, silence) and explicit (cancellation, repetition) signals as noisy indicators of misclassification, enabling online pattern updates without manual annotation. On a proprietary dataset of 3,667 interactions, ASCIL reduces errors by up to 54.27% relative on a session‑disjoint subset and 24.39% at a 0.90 threshold, while adding less than 60 ms of latency and improving intentional acceptance rates.

By Preeti Saraswat, Divya Neelagiri, Anil Yadav
arXiv AI
Sep 2

Conversation Coach: A Voice-enabled AI System that Helps Practice Difficult Workplace Conversations

Conversation Coach is a voice‑first AI system that lets managers rehearse difficult workplace conversations in a realistic spoken format. It tackles low‑latency interaction, adaptive bot personalities that simulate various employee types, and personalized feedback on content and policy compliance. The authors compare an end‑to‑end speech‑to‑speech model with a cascaded approach, finding the former offers lower latency and cost, while the cascaded model provides better reasoning for coaching quality, and they deployed the cascaded architecture to 40,000+ managers over six months.

By Fanyou Wu, Suraj Maharjan, Ainur Yessenalina, Dennis Xu Chen, Rahul Srivastava, Srinivasan H. Sengamedu