arXiv Machine Learning

Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

The paper introduces a scalable method for creating pseudo-triplet data—comprising a reference utterance, a direction text, and a modified utterance—to train direction‑following text‑to‑speech (TTS) systems. It uses an impression‑controllable TTS model to generate style variations and a large language model to generate natural language directions from estimated impression differences. Experiments show that these pseudo‑triplets enable stable speaker‑preserving modifications, and combining them with recorded data further improves direction alignment while maintaining speaker similarity.

arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen
arXiv Machine Learning
Aug 17

VoiceDesigner: Text-to-Voice Generation and Editing via Unified Diffusion Modeling and Data Augmentation

arXiv:2608. 13613v1 Announce Type: cross Abstract: Recent breakthroughs in generative models have made text-to-voice generation (TTV) possible, enabling the synthesis of speech directly from textual voice descriptions.

By Jiarui Hai, Karan Thakkar, Ke Chen, Yunyun Wang, Jiaqi Su, Rithesh Kumar, Mounya Elhilali, Zeyu Jin
arXiv AI
Aug 24

Poly-InstructTTS: Learning In-the-Wild Expressive Speech Synthesis from Open-Ended Instructions

Poly-InstructTTS is a text‑to‑speech system that learns expressive speech from open‑ended natural‑language instructions using a 1,000‑hour, 1,000‑plus emotion and style annotated audiovisual corpus. The approach employs a prompt‑free GPT with attribute‑based thinking tokens and a flow‑matching module to inject timbre from reference audio, and includes a speaker fine‑tuning procedure to transfer instruction control while preserving speaker persona. Experiments demonstrate strong instruction adherence and expressiveness, with audio demos and an expanded test set available on the project page.

By Junhui Zhang, Qianhui Xu, Qingxiang Guo, Dawei Yang, Ling Miao, Qiangqiang Wang, Yang Song
arXiv AI
Jun 24

ZONOS2 Technical Report

arXiv:2606. 24320v1 Announce Type: cross Abstract: We present ZONOS2 8B, our latest TTS model, which achieves state-of-the-art naturalness, prosody, and voice cloning fidelity.

By Gabriel Clark, Sofian Mejjoute, Mohamed Osman, George Close, Beren Millidge