BaltiVoice: A Speech Corpus and Fine-tuned Whisper ASR System for the Balti Language
arXiv:2606. 03504v1 Announce Type: cross Abstract: We present BaltiVoice, a 16.
We present BaltiVoice, a 16. 8-hour read-speech corpus for Balti (ISO 639-3: bft), a Tibetic language spoken in Gilgit-Baltistan, Pakistan, with no prior publicly available ASR resources.
arXiv:2606. 03504v1 Announce Type: cross Abstract: We present BaltiVoice, a 16.
The study fine‑tunes the Whisper Small model for Automatic Speech Recognition (ASR) in Baniwa, an indigenous Arawakan language. Using a 0.54‑hour corpus of 1,373 manually transcribed recordings, the fine‑tuned model achieved a Word Error Rate of 37.5% and a Character Error Rate of 7.45%. These results provide an initial baseline for Baniwa ASR and suggest that multilingual foundation models can be adapted to extremely low‑resource languages.
arXiv:2607. 17164v1 Announce Type: new Abstract: Developing Automatic Speech Recognition (ASR) for morphologically rich, low-resource languages such as Assamese is challenging due to insufficient annotated speech data.
The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.
arXiv:2609.14542v1 Announce Type: new Abstract: Neyshekar is presented as an open Persian read-speech corpus designed for coverage of both formal and informal language, named entities, and longer utt...
arXiv:2608. 12327v1 Announce Type: cross Abstract: Multilingual pretrained models nominally support Nepali, yet no controlled benchmark has compared them under a single fine-tuning protocol.
arXiv:2607. 23808v1 Announce Type: cross Abstract: In this work, we introduce Indic DiarBench, a speaker diarization and ASR benchmark dataset spanning all 22 scheduled languages of India.
arXiv:2603. 04710v2 Announce Type: replace-cross Abstract: Recent advances in automatic speech recognition (ASR) and speech enhancement have strengthened the common belief that cleaner audio should lead to more accurate transcription.
arXiv:2609.24199v1 Announce Type: new Abstract: Evaluation benchmarks for Indian language automatic speech recognition (ASR) suffer from two systematic biases: optimistic scores from clean, controlle...
arXiv:2606. 07608v1 Announce Type: cross Abstract: We present a systematic study of fine-tuning OpenAI's Whisper large-v3 for Swiss German ASR, using 1,367 hours of broadcast speech paired with Standard German subtitles as weak supervision.
The TutlAit v1 dataset is a crowdsourced corpus of Moroccan Tamazight speech paired with Modern Standard Arabic transcriptions and explicit regional accent labels. It contains 13,384 audio files (≈20.9 hours) collected via a web application, with volunteers contributing through text‑to‑audio and audio‑to‑text workflows, and includes additional segments from freely available media. The dataset covers Atlas, Souss, Rif, and Kabyle varieties, making it a valuable resource for speech recognition, translation, and accent identification in an under‑resourced language.
SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.