Making automatic speech recognition work on large files with Wav2Vec2 in ๐ค Transformers
Related stories
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐ค Transformers
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with ๐ค Transformers
Beyond Short Segments : Expanding Speaker Embeddings with Vector Archives
arXiv:2609.25007v1 Announce Type: cross Abstract: The performance of state-of-the-art speaker verification (SV) systems severely degrades on short utterances due to insufficient speaker-specific info...
A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.
P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
arXiv:2609.24138v1 Announce Type: cross Abstract: Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has c...
Enabling automatic transcription of child-centered audio recordings from real-world environments
The paper presents a method to automatically identify utterances in child-centered daylong audio recordings that can be reliably transcribed by modern ASR systems, enabling accurate transcription of a substantial portion of the speech. On four English corpora, the approach achieves a median WER of 0% and a mean WER of 16% when transcribing 30% of the total speech, compared to a median WER of 52% when transcribing all speech. Word frequency distributions from the automatic transcripts correlate strongly with manual annotations (rโฏ=โฏ0.94 overall, rโฏ=โฏ0.99 for frequent words).
Audio-Native Speech Recognition with a Frozen Discrete-Diffusion Language Model
arXiv:2607. 13013v1 Announce Type: new Abstract: Automatic speech recognition is dominated by autoregressive decoders that emit one token at a time.
P2Flow: Phoneme-aware Progressive Flow Matching for Extreme Speech Super-Resolution
Generative models have recently demonstrated considerable promise in speech super-resolution (SSR). Nevertheless, the majority of existing work has concentrated on standard or versatile SSR configurat...
VOSSA: Voiceprint Optimization for Streaming Speech Architectures
arXiv:2609.38887v1 Announce Type: cross Abstract: Real-time voice conversion (VC) systems commonly rely on pretrained speaker embeddings from automatic speaker verification (ASV) models. While effect...
Advancing Speaker-Based Vocal Effort Classification with WavLM and Data Augmentation in Naturalistic Non-Calibrated Speech Recordings
arXiv:2606. 27543v1 Announce Type: cross Abstract: The variations in vocal effort range (e.
Phonetic Error Analysis of Raw Waveform Acoustic Models
arXiv:2606. 07030v1 Announce Type: cross Abstract: We analyse error patterns of raw waveform acoustic models on TIMIT phone recognition beyond the overall phone error rate (PER).