arXiv Computation and Language By Eunjung Yeo, Kwanghee Choi, Krupaben Kothadia, Visar Berisha, Julie M. Liss, David R. Mortensen, David Harwath

Quantifying Consonant Contributions to Word Intelligibility via Acoustic Masking

Read the original on arXiv Computation and Language →

The study introduces a scalable acoustic‑masking method to quantify how much each consonant contributes to word intelligibility. By silencing individual consonants in isolated words and measuring misrecognition rates with three ASR models, the authors define a mask‑induced misrecognition rate (MMR). Across English, Spanish, German, and Czech, MMR negatively correlates with phoneme frequency and positively with functional load, revealing that consonant importance varies by language.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
6d ago

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

The paper introduces a method that allows automatic speech recognition systems to learn new words during test time using unlabeled data. It combines a frozen CTC acoustic model for spellings, a frozen language model for detecting out‑of‑vocabulary words, and an adaptation module that expands the vocabulary by learning lexical token representations from CTC-generated candidates. Experiments on LibriSpeech and dysarthric speech data show relative character‑error‑rate reductions of up to 14.97% and 6.67% for recurring OOV words, respectively.

By Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo