Making automatic speech recognition work on large files with Wav2Vec2 in ๐ค Transformers
Related stories
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐ค Transformers
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with ๐ค Transformers
A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
arXiv:2607. 13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time.
Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings
arXiv:2606. 27543v1 Announce Type: cross Abstract: The variations in vocal effort range (e.
Phonetic Error Analysis of Raw Waveform Acoustic Models
arXiv:2606. 07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER).
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time. We ask whether a discrete diffusion language model can transcribe speech instead, refining a whole transcript in parallel over a small number of denoising steps.
An Empirical Recipe for Universal Phone Recognition
arXiv:2603. 29042v2 Announce Type: replace-cross Abstract: Phone recognition (PR) is a key enabler of multilingual and low-resource speech processing tasks, yet robust performance remains elusive.
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
arXiv:2608. 10206v1 Announce Type: new Abstract: Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech.
Towards Audio Token Compression in Large Audio Language Models
arXiv:2511. 20973v2 Announce Type: replace-cross Abstract: Large Audio Language Models (LALMs) deliver strong performance across speech and audio tasks, but their audio encoders generate high-rate token sequences (e.
Edge Phoneme Recognition for Children's Speech through Age-Aware Training
Detecting phonemes from children's speech has historically been difficult due to the scarcity of training data, and unique characteristics of children's speech. During a phoneme detection competition, we found that training a lightweight model to predict the age of the learner, as well as the phoneme sequence, enabled a 94M-parameter model to outperform WavLM Large models (317M) on the target DrivenData distribution, and fall within approximately 0.