Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with ๐ค Transformers
Related stories
Fine-Tune Wav2Vec2 for English ASR in Hugging Face with ๐ค Transformers
Boosting Wav2Vec2 with n-grams in ๐ค Transformers
Making automatic speech recognition work on large files with Wav2Vec2 in ๐ค Transformers
Fine-Tune Whisper For Multilingual ASR with ๐ค Transformers
Fine-Tune MMS Adapter Models for low-resource ASR
Contrastive Regularization for Accent-Robust ASR
arXiv:2605. 03297v2 Announce Type: replace-cross Abstract: ASR systems based on self-supervised acoustic pretraining and CTC fine-tuning achieve strong performance on native speech but remain sensitive to accent variability.
Comparative Analysis of Multilingual Pre-trained Models for Nepali Automatic Speech Recognition
arXiv:2608. 12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol.
Data Scale, Not Latency, Shapes Cross-Lingual Encoder Transfer in Streaming ASR
arXiv:2606. 24169v1 Announce Type: new Abstract: Adapting a streaming speech recognition model to a new language requires choosing between two plausible warm starts: a multilingual (ML) encoder or an English-only (EN) encoder.
Benchmarking Language Modeling for Lossless Compression of Full-Fidelity Audio
arXiv:2603. 08683v2 Announce Type: replace-cross Abstract: Autoregressive "language" models (LMs) trained on raw waveforms can be repurposed for lossless audio compression, but prior work is limited to 8-bit audio, leaving open whether such approaches work for practical settings (16/24-bit) and can compete with existing codecs.
A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition
arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.
The Equalizer: Introducing Shape-Gain Decomposition in Neural Audio Codecs
arXiv:2602. 15491v2 Announce Type: replace-cross Abstract: Neural audio codecs (NACs) typically encode the short-term energy (gain) and normalized structure (shape) of speech/audio signals jointly within the same latent space.