Curriculum-Based Noise Adaptation for Phoneme-to-Text Reconstruction in Visual Speech Recognition
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
arXiv:2606.00751v2 Announce Type: replace Abstract: Visual Speech Recognition (VSR) aims to recognize speech from visual cues such as lip movements. Still, its performance is fundamentally limited by...
arXiv:2604.06487v2 Announce Type: replace Abstract: Conventional end-to-end automatic speech recognition (ASR) systems rely on paired speech-text data for domain adaptation. Recent LLM-based ASR arch...
arXiv:2509.10452v3 Announce Type: replace-cross Abstract: Pretrained automatic speech recognition (ASR) models such as Whisper perform well but still need domain adaptation to handle unseen parlance....
The paper introduces a phoneme-guided text-to-speech (TTS) augmentation pipeline for automatic speech recognition (ASR) that links multilingual speech generation with candidate-text selection and reference-speech quality control. It proposes phoneme-frequency-guided selection (PFGS), which prioritizes candidate texts containing common phonetic content based on real ASR training transcripts. Experiments across four languages and 13 test sets show that random text selection improves recognition on 11 test sets, while PFGS further improves nine test sets with relative word error rate reductions up to 19.3%, and reference-speech filtering also contributes to performance gains.
arXiv:2609.09757v1 Announce Type: cross Abstract: Real-time MRI (rtMRI) captures the dynamics of the entire vocal tract during speech, but labeled data are scarce and the modality - single-slice, gra...
arXiv:2609.10394v1 Announce Type: cross Abstract: Current audio-visual speech recognition (AVSR) benchmarks, like LRS3, rely heavily on clean, scripted and rehearsed speech. They fail to reflect the...