arXiv Machine Learning

SoniSpeech: A Large-Scale Open-Vocabulary Tri-Modal Dataset for Wearable Silent Speech Interfaces

arXiv:2608. 00803v1 Announce Type: cross Abstract: Wearable silent speech interfaces (SSIs) are limited to small, closed vocabularies.

arXiv Computation and Language
Sep 25

YODAS v3: Over 1 Million Hours of High-Bandwidth, Stereophonic, Multilingual Speech

YODAS v3 is a weakly‑labeled speech corpus that offers more than 1.1 million hours of 48 kHz multi‑channel audio across 147 languages, making it the largest open speech dataset available and the first large‑scale collection with high‑fidelity stereo audio. The authors detail a new collection methodology that balances language representation, achieving 22 languages with over 10 k hours and 73 languages with over 5 k hours of data. They also analyze language, audio, and transcription quality, and demonstrate the dataset’s utility by training baseline speech‑recognition and neural‑codec models.

By William Chen, Shinnosuke Takamichi, Sayaka Shiota, Satoru Fukayama, Samuele Cornell, Shinji Watanabe
arXiv Computation and Language
Sep 25

DiscoPhon: Benchmarking the Unsupervised Discovery of Phoneme Inventories With Discrete Speech Units

DiscoPhon is a multilingual benchmark designed to evaluate unsupervised phoneme discovery from discrete speech units. It includes 6 development and 6 test languages that cover a wide range of phonemic contrasts, and requires systems to generate discrete units mapped to a predefined phoneme inventory using only 10 hours of speech from an unseen language. The benchmark assesses unit quality, recognition, and segmentation, and provides four pretrained multilingual HuBERT and SpidR baselines that demonstrate current models can produce units that correlate well with phonemes, though performance varies across languages.

By Maxime Poli, Manel Khentout, Angelo Ortiz Tandazo, Ewan Dunbar, Emmanuel Chemla, Emmanuel Dupoux
arXiv Machine Learning
Sep 18

Enabling automatic transcription of child-centered audio recordings from real-world environments

The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (r = 0.94 overall, r = 0.99 for frequent words).

By Daniil Kocharov, Azarias Galama, Okko R\"as\"anen
arXiv Machine Learning
Jun 8

SEAM: Shortcut-Aware Real-Time Detection of Scripted vs. Spontaneous Speech for Interview Guardrails

arXiv:2606. 06837v1 Announce Type: cross Abstract: Scripted vs spontaneous speech detection is appealing for interview guardrails, but benchmark performance can be inflated by shortcuts tied to corpus identity, channel conditions, and recording artifacts rather than speaking style itself.

By Vsevolod (V.), Kovalev, Pranay Manocha