P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
Read the original on Hugging Face Trending Papers →The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
The Flow has not summarised this story yet — read it at Hugging Face Trending Papers.
arXiv:2609.24138v1 Announce Type: cross Abstract: Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has c...
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
PitchFlower is a flow‑based neural audio codec that offers explicit pitch controllability by flattening and randomly shifting F0 contours during training while conditioning on the true F0 to reconstruct the original audio. A vector‑quantization bottleneck blocks pitch recovery, and a flow‑based decoder produces high‑quality audio. Experiments demonstrate that PitchFlower matches DSP baselines in pitch accuracy, surpasses state‑of‑the‑art neural codecs in audio quality, and remains robust even when trained on WORLD‑transformed audio, effectively removing vocoder artifacts.
arXiv:2607. 00363v1 Announce Type: cross Abstract: Flow Matching (FM) has emerged as a powerful paradigm for speech generation but remains constrained by high inference latency and timbre leakage.
arXiv:2412. 11449v2 Announce Type: replace-cross Abstract: We propose WHISPER-GPT: A generative large language model (LLM) for speech and music that allows us to work with continuous audio representations and discrete tokens simultaneously as part of a single architecture.
Conditional Flow Matching models for text‑to‑speech often produce incoherent frequency evolution during inference. The authors propose a training‑free, frequency‑selective boosting strategy that uses the Discrete Wavelet Transform to dynamically modulate mel‑spectrogram sub‑bands during ODE integration, penalizing aggressive low‑frequency growth while boosting lagging high‑frequency details. Across multiple architectures, this method reduces the number of function evaluations from 32 to 26 and improves Frechet Audio Distance by up to 61% without harming mean opinion scores, speaker similarity, or intelligibility.