arXiv Machine Learning
1d ago

Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction

The paper introduces a scalable method for creating pseudo-triplet data—comprising a reference utterance, a direction text, and a modified utterance—to train direction‑following text‑to‑speech (TTS) systems. It uses an impression‑controllable TTS model to generate style variations and a large language model to generate natural language directions from estimated impression differences. Experiments show that these pseudo‑triplets enable stable speaker‑preserving modifications, and combining them with recorded data further improves direction alignment while maintaining speaker similarity.

By Kenichi Fujita, Yusuke Ijima
arXiv AI
Jun 18

Reference-Driven Multi-Speaker Audio Scene Generation from In-the-Wild Priors

arXiv:2606. 19325v1 Announce Type: cross Abstract: Existing multi-speaker dialogue systems bind speakers to utterances through structured supervision: per-turn tags, multi-stream transcriptions, or learnable speaker embeddings.

By Michael Finkelson, Daniel Segal, Eitan Richardson, Shahar Armon, Nani Goldring, Poriya Panet, Nir Zabari, Benjamin Brazowski, Or Patashnik, Yoav HaCohen