arXiv AI

Multi-Dimensional Prosody Judgment For Live Streaming Speech Synthesis

The paper introduces Live-ProsodyJudge (LPJ), a cost‑effective pairwise evaluator distilled from Gemini for assessing fine‑grained prosody in live streaming TTS. It identifies a flaw called verdict coupling, where multi‑dimensional scores collapse into a single preference, and proposes Decoupled‑Live‑ProsodyJudge (D‑LPJ) to eliminate this issue through masking and a span‑local GRPO strategy. Experiments show LPJ outperforms a single Gemini call in accuracy, and D‑LPJ provides independent dimension judgments, achieving high alignment with human top‑3 selections in a TTS candidate tournament.

arXiv Machine Learning
Sep 15

The Limits of Reference-Free Speech Quality Metrics as Evaluators and Rewards on Modern Text-to-Speech

arXiv:2609.13150v1 Announce Type: cross Abstract: Reference-free quality predictors such as UTMOS, DNSMOS and SCOREQ are the de facto automatic evaluators for text-to-speech (TTS) and are increasingl...

By Antonis Asonitis, Juan Pablo Zuluaga Gomez, Francesco Verdini, Aref Farhadipour, Marzieh Razavi, Pierre-Edouard Honnet, Vijeta Avijeet
arXiv AI
Aug 11

Beyond Naturalness: Probing Automated Text-To-Speech Evaluators on Linguistically Grounded Dimensions

arXiv:2608. 09930v1 Announce Type: cross Abstract: Automated Text-to-Speech (TTS) evaluation methods (Mean Opinion Score (MOS) predictors and Audio Large Language Models (Audio-LLM) judges) are expected to reflect human perception, yet it is unclear how well they capture the distinct aspects of speech that listeners actually perceive.

By Oluwanifemi Bamgbose, Simon Rosen, Jash Shah, Lindsay Devon Brin, Hoang H Nguyen, Anke Koelzer, Rachel Hansen, Tara Bogavelli, Fanny Riols