arXiv AI

Human-1 by Josh Talks: A Full-Duplex Conversational Modeling Framework in Hindi using Real-World Conversations

The paper introduces Human-1, the first open and reproducible full‑duplex spoken dialogue system for Hindi. It adapts the Moshi architecture with a custom Hindi tokenizer and trains on 26,000 hours of real spontaneous conversations from 14,695 speakers, enabling the model to learn natural turn‑taking, interruptions, overlaps, and backchannels. A two‑stage training process—large‑scale pre‑training followed by fine‑tuning on 1,000 hours—yields a system that, according to both automatic metrics and human judgments, generates natural and meaningful full‑duplex conversational behavior in Hindi.

arXiv Computation and Language
3d ago

Learning Natural Conversational Behavior in Tandem Speech-to-Speech Models with Randomized Guidance

The paper introduces a method called randomized intermediate guidance for training tandem speech-to-speech models, where a large language model (LLM) acts as a backend providing candidate responses while the user is speaking. Instead of simulating the backend’s guidance, the approach derives guidance directly from the conversation corpus, using target responses for informative guidance and randomly sampled responses to simulate irrelevant updates. Experiments on synthetic dialogues and 3.8k hours of real conversations show that this technique yields response quality comparable to LLM-generated baselines while improving natural turn‑taking and audio‑judge naturalness.

By Manato Yaguchi, Yotaro Kubo, Hikaru Asano, So Kuroki
arXiv AI
Aug 19

Multi-turn Conversational AI from Text to Multimodal Interaction: Data, Models, Evaluation, and Open Challenges

The article surveys multi‑turn conversational AI, highlighting its shift from isolated text prompts to sustained, multimodal interactions that involve clarifying goals, revising requests, and switching topics. It reviews literature across text‑only dialogue, AudioLLMs, multimodal and omni‑modal systems, and tool‑augmented agents, organizing findings around datasets, models, training, evaluation, and cross‑cutting challenges. The analysis reveals that while multimodal perception and action have progressed rapidly, systems still struggle with persistent memory, cross‑turn grounding, full‑duplex interaction, robust evaluation, and cultural alignment.

By Syeda Faiza Ahmed, Zien Sheikh Ali, Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv AI
Sep 24

Training Intelligent Voice Assistant Wakeup with Controllable Synthetic Conversations

The paper presents a new wake‑up system for voice assistants that goes beyond simple keyword spotting by adding contextual trigger detection. After the wake word is heard, the system reasons to differentiate between actual user commands and unrelated speech, enabling more efficient and context‑aware interactions. A data‑generation architecture is introduced that creates a 62.3‑hour corpus of controllable multi‑speaker conversations, including direct invocations, contextual follow‑ups, and non‑addressed speech, and experimental results confirm the approach’s effectiveness across varied synthetic scenarios.

By Marcin Sowa\'nski, Kacper Leszczy\'nski, Kacper Krzywicki, Krzysztof Wodnicki
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.

arXiv Computation and Language
Sep 14

DuplexDrama: A Synthesized Dialogue Dataset with Scenarios, Full-Duplex Behaviors, Expressive Speech, and Sound Events

DuplexDrama is a newly announced synthesized spoken dialogue dataset that uniquely combines complete persona and scenario settings, three full‑duplex behaviors (interruption, backchannel, incomplete), expressive speech with persona‑aligned emotion labels, and script‑aware sound events. The dataset was created through a four‑stage pipeline and validated for quality on both scripts and audio, yielding over 2,000 hours of audio featuring 64 voices across 13 personas and 5 age groups, with 3.8% of turns containing full‑duplex behaviors. A curated bilingual subset of 6,400 dialogues (800 hours total) will be released to support research in full‑duplex spoken dialogue models, and evaluation prompts will accompany the dataset.

By Qingxiang Guo, Wenke Fan, Shuofeng Zhao, Dawei Yang, Zhiyang Zhou, Yingxin Shang, Hongwei Cai, Zhou Wang, Weixu Wang, Lin Yang, Shuran Zhou, Yang Song