arXiv Computation and Language

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

YODAS v3 is a weakly‑labeled speech corpus that offers more than 1.1 million hours of 48 kHz multi‑channel audio across 147 languages, making it the largest open speech dataset available and the first large‑scale collection with high‑fidelity stereo audio. The authors detail a new collection methodology that balances language representation, achieving 22 languages with over 10 k hours and 73 languages with over 5 k hours of data. They also analyze language, audio, and transcription quality, and demonstrate the dataset’s utility by training baseline speech‑recognition and neural‑codec models.

arXiv Computation and Language
Sep 15

CVSS-X: A Multilingual Speech-to-Speech Translation Corpus for 28 Languages

arXiv:2609.13413v1 Announce Type: new Abstract: We introduce CVSS-X, a large-scale synthetic speech-to-speech translation corpus that extends CVSS by reversing the translation direction. While CVSS t...

By Lucas Rafael Stefanel Gris, Alef Iury Siqueira Ferreira, Frederico Santos de Oliveira, Augusto Seben da Rosa, Alexandre Costa Ferro Filho, Arlindo Rodrigues Galv\~ao Filho, Anderson da Silva Soares
arXiv AI
Jun 10

Linguistically Augmented Audio Speech Data (LinguAS)

arXiv:2606. 10246v1 Announce Type: cross Abstract: Maliciously-created fake speech, including deepfaked and spoofed audio, is proliferating at an alarming rate, and detection models are racing to stay ahead of the curve.

By Ashley R. Keaton, Zahra Khanjani, Christine Mallinson, Vandana P. Janeja
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
Hugging Face Trending Papers
Jun 2

Efficient ASR Training with Conversations that Never Happened

Conversational ASR for lower-resource languages and niche domains is limited by the scarcity of domain-matched multi-speaker training data. We propose an augmentation pipeline that generates scenario-level dialogues with participant metadata, maps speaker attributes to TTS voice profiles, and assembles synthesized utterances into speaker-aware simulated conversations.