arXiv:2609.09719v1 Announce Type: new
Abstract: Text-aligned speech tokenization methods have emerged to better align speech tokens with LLM token spaces, enabling more effective utilization of pretr...
By Kang-wook Kim, Jinyoung Park, Jinsoo Kim, Sehun Lee, Sang Hoon Woo, Gunhee Kim
arXiv:2608. 01281v1 Announce Type: cross Abstract: Phoneme-based multilingual automatic speech recognition (ASR) can share acoustic evidence across languages more directly than language-specific subword modeling.
By Saierdaer Yusuyin, Nanling Jiang, Hao Huang, Zhijian Ou
The paper introduces a method that allows automatic speech recognition systems to learn new words during test time using unlabeled data. It combines a frozen CTC acoustic model for spellings, a frozen language model for detecting out‑of‑vocabulary words, and an adaptation module that expands the vocabulary by learning lexical token representations from CTC-generated candidates. Experiments on LibriSpeech and dysarthric speech data show relative character‑error‑rate reductions of up to 14.97% and 6.67% for recurring OOV words, respectively.
By Mengqi Wang, Mark A. Hasegawa-Johnson, Haolong Zheng, Chang D. Yoo
arXiv:2607. 06831v1 Announce Type: cross Abstract: Speech-to-text alignment means finding the temporal boundaries of each word in the audio.
By Albert Zeyer, Ralf Schl\"uter, Hermann Ney
arXiv:2606. 27698v1 Announce Type: cross Abstract: Automatic Speech Recognition systems are notoriously both sensitive to adversarial and benign perturbations.
By Andrew C. Cullen, Neil Marchant, Jiani Xie, Paul Montague, Benjamin I. P. Rubinstein
The paper introduces a training‑free speech‑and‑text‑to‑pronunciation (ST2P) pipeline that combines lexical candidates from G2P tools with acoustic rescoring using frozen pretrained S2P models. By performing a left‑to‑right greedy search over whole‑sequence negative log‑likelihoods, the method achieves a dramatic reduction in character error rate on Japanese corpora, outperforming both baseline G2P/S2P approaches and commercial multimodal LLMs. The approach is also significantly faster—3–3.5× faster than beam search and twice as fast as direct decoding—while maintaining high accuracy across multiple languages.
By Hikaru Asano, Yotaro Kubo, So Kuroki