Scalable Direction-Following TTS via Voice Impression-Guided Pseudo Triplet Construction
Read the original on arXiv Machine Learning →The paper introduces a scalable method for creating pseudo-triplet data—comprising a reference utterance, a direction text, and a modified utterance—to train direction‑following text‑to‑speech (TTS) systems. It uses an impression‑controllable TTS model to generate style variations and a large language model to generate natural language directions from estimated impression differences. Experiments show that these pseudo‑triplets enable stable speaker‑preserving modifications, and combining them with recorded data further improves direction alignment while maintaining speaker similarity.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Machine Learning.