arXiv Machine Learning

P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution

arXiv Machine Learning
Jul 20

RobustSpeechFlow: Learning Robust Text-to-Speech Trajectories via Augmentation-based Contrastive Flow Matching

arXiv:2605. 22083v2 Announce Type: replace-cross Abstract: While flow-matching text-to-speech (TTS) achieves strong zero-shot speaker similarity and naturalness, it remains susceptible to content fidelity issues, particularly skip and repeat errors from imperfect alignment.

By Jinhyeok Yang, Hyeongju Kim, Yechan Yu, Joon Byun, Frederik Bous, Juheon Lee
arXiv Machine Learning
1d ago

Harmonizing Spectral Evolution in Conditional Flow Matching for TTS

Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.

By Isha Pandey, Varad Deshpande, Abhijat Bharadwaj, Ganesh Ramakrishnan
arXiv Machine Learning
Sep 11

PitchFlower: A flow-based neural audio codec with pitch controllability

PitchFlower is a flow‑based neural audio codec that offers explicit pitch controllability by flattening and randomly shifting F0 contours during training while conditioning on the true F0 to reconstruct the original audio. A vector‑quantization bottleneck blocks pitch recovery, and a flow‑based decoder produces high‑quality audio. Experiments demonstrate that PitchFlower matches DSP baselines in pitch accuracy, surpasses state‑of‑the‑art neural codecs in audio quality, and remains robust even when trained on WORLD‑transformed audio, effectively removing vocoder artifacts.

By Diego Torres, Axel Roebel, Nicolas Obin
arXiv Computation and Language
Sep 18

Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation

The paper introduces a phoneme-guided text-to-speech (TTS) augmentation pipeline for automatic speech recognition (ASR) that links multilingual speech generation with candidate-text selection and reference-speech quality control. It proposes phoneme-frequency-guided selection (PFGS), which prioritizes candidate texts containing common phonetic content based on real ASR training transcripts. Experiments across four languages and 13 test sets show that random text selection improves recognition on 11 test sets, while PFGS further improves nine test sets with relative word error rate reductions up to 19.3%, and reference-speech filtering also contributes to performance gains.

By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang, Wei Xu