BareWave: Waveform-Native Flow-Matching Text-to-Speech
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
arXiv:2606. 14791v1 Announce Type: cross Abstract: Self-supervised learning advances audio representation for multimedia analysis.
arXiv:2606. 09048v1 Announce Type: cross Abstract: Removing intermediate representations and separately trained decoding stages has become an important direction in generative modeling.
arXiv:2606. 30700v1 Announce Type: cross Abstract: Self-supervised learning enables audio representations that transfer across domains and tasks.
The paper investigates how procedural audio should be scaled for effective pre‑training and whether training strategies from natural audio transfer to procedural data. By separating scale into formula‑class coverage and within‑class rendering diversity, the authors show that each type of scale benefits different learning formulations and downstream tasks. Their experiments reveal that procedural audio prefers lower mask ratios, and that it exhibits lower patch diversity and stronger temporal predictability compared to natural audio, leading to a proposal for source‑aware procedural pre‑training.
arXiv:2606. 06357v1 Announce Type: cross Abstract: Continuous audio autoencoders reconstruct waveforms well but often produce latents with weak structure for understanding, while self-supervised audio encoders capture semantics but are not directly decodable.
arXiv:2606. 30356v1 Announce Type: cross Abstract: We propose Online Latent prediction with Invariant Views and rEconstruction (OLIVE), a self-supervised speech representation learning framework that jointly optimizes analysis and synthesis objectives.
SCAPES is a lightweight, resource‑efficient generative model that synthesizes high‑fidelity environmental sounds with high‑level semantic control. It operates on the continuous latent manifold of a neural audio codec, using a segmentation strategy and a Continuous Normalizing Flow to model latent trajectories. A 36‑million‑parameter instance can be trained on limited, uncurated data with a single consumer‑grade GPU, achieving convergence in roughly twice the source audio duration and enabling smooth semantic interpolation.
TUTTI is a new pre‑training framework for audio‑to‑score transcription that uses a large, fully synthetic multi‑instrument dataset generated by a symbolic music model. The approach trains a standard Transformer encoder‑decoder on these synthetic audio‑score pairs, producing a stronger foundational representation than single‑instrument training. When fine‑tuned on real datasets, TUTTI surpasses prior methods, achieving state‑of‑the‑art results and demonstrating strong cross‑instrument transferability.
arXiv:2607. 08545v1 Announce Type: cross Abstract: End-to-end neural audio models achieve high-fidelity compression and generation.
arXiv:2608. 19863v1 Announce Type: cross Abstract: Self-supervised learning (SSL) has driven substantial progress in audio representation learning, though existing methods have increasingly relied on elaborate pre-training recipes to reach competitive performance.
Generalizable Audio-to-Score (A2S) transcription is fundamentally constrained by the severe scarcity of high-quality, real-world paired data. Relying solely on existing human-annotated datasets often...
arXiv:2606. 29031v1 Announce Type: cross Abstract: In regulated domains such as banking and healthcare, where privacy constraints make real speech costly to collect and retain, synthetic speech from modern text-to-speech (TTS) is an appealing alternative for training automatic speech recognition (ASR) without exposing sensitive customer recordings.
GrainSpeech is a compact speech synthesis model that uses a fixed‑receptive‑field convolutional encoder to reduce pitch, energy, and duration prediction errors by 36.0%, 17.3%, and 3.4% respectively. It introduces a Mel‑specific gradient‑variance supervision that improves fine‑scale variation while avoiding quality degradation. With only 264.8K parameters, GrainSpeech achieves 17.9× real‑time Mel generation on a microcontroller and attains UTMOS scores comparable to much larger models, using less than 1.5% of their parameters.