Fine-Tune Whisper For Multilingual ASR with 🤗 Transformers
Related stories
BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such...
Fine-Tune XLSR-Wav2Vec2 for low-resource ASR with 🤗 Transformers
BuzzASR: A Swarm of 100+ Monolingual Speech Recognition Models
arXiv:2609.09554v1 Announce Type: new Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. L...
Fine-Tune W2V2-Bert for low-resource ASR with 🤗 Transformers
On the Interpretability of Whisper Encodings Using Sparse Autoencoders
arXiv:2605.12225v3 Announce Type: replace Abstract: While deep transformer-based models have advanced rapidly, their internal mechanisms remain largely a mystery. Recent work has prioritized understa...
How to generate text: using different decoding methods for language generation with Transformers
Sequential Adapter Stacking for Cross-Lingual Low-Resource ASR
arXiv:2609.15758v1 Announce Type: new Abstract: Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed...
BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language
arXiv:2606. 03504v1 Announce Type: cross Abstract: We present BaltiVoice, a 16.
BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language
We present BaltiVoice, a 16. 8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources.
Phoneme-guided TTS augmentation for ASR: A unified pipeline and multilingual evaluation
The paper introduces a phoneme-guided text-to-speech (TTS) augmentation pipeline for automatic speech recognition (ASR) that links multilingual speech generation with candidate-text selection and reference-speech quality control. It proposes phoneme-frequency-guided selection (PFGS), which prioritizes candidate texts containing common phonetic content based on real ASR training transcripts. Experiments across four languages and 13 test sets show that random text selection improves recognition on 11 test sets, while PFGS further improves nine test sets with relative word error rate reductions up to 19.3%, and reference-speech filtering also contributes to performance gains.
Automatic Speech Recognition for Multilingual Oral History Research
The paper examines how Automatic Speech Recognition (ASR) tools, particularly Whisper, are being used in community-led heritage language preservation, focusing on Cantonese oral histories in New Zealand. It reports that the best Whisper configuration achieved a 12.10 % Word Error Rate (WER) but struggled with non‑English segments, yet it can produce a first‑pass transcription in only 1 % of the time required for manual transcription.