YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech
Read the original on arXiv Computation and Language →YODAS v3 is a weakly‑labeled speech corpus that offers more than 1.1 million hours of 48 kHz multi‑channel audio across 147 languages, making it the largest open speech dataset available and the first large‑scale collection with high‑fidelity stereo audio. The authors detail a new collection methodology that balances language representation, achieving 22 languages with over 10 k hours and 73 languages with over 5 k hours of data. They also analyze language, audio, and transcription quality, and demonstrate the dataset’s utility by training baseline speech‑recognition and neural‑codec models.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.