arXiv Computation and Language By Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo

Learning New Words from Unlabeled Test Data in Automatic Speech Recognition

Read the original on arXiv Computation and Language →

The paper introduces a method that allows automatic speech recognition systems to learn new words during test time using unlabeled data. It combines a frozen CTC acoustic model for spellings, a frozen language model for detecting out‑of‑vocabulary words, and an adaptation module that expands the vocabulary by learning lexical token representations from CTC-generated candidates. Experiments on LibriSpeech and dysarthric speech data show relative character‑error‑rate reductions of up to 14.97% and 6.67% for recurring OOV words, respectively.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
5d ago

A Training Criterion with Token-Level Tolerance to Transcription Ambiguity for Automatic Speech Recognition

The paper introduces a token‑level extension of Omni‑Temporal Classification (OTC) for automatic speech recognition, allowing unsupported tokens to be bypassed while preserving supervision for the rest of the word. Across 19 languages and three corpora, this token‑level OTC consistently outperforms standard CTC, achieving the lowest mean word error rate on every dataset and a 9.45% average relative WER reduction. A predictive‑entropy‑indexed schedule replaces epoch‑based relaxation, reducing training‑length dependence while maintaining performance.

By Saurabh Kumar, Diptiman Mohanta, Prasanta Kumar Ghosh