arXiv Computation and Language By Biraj Subedi

SpeakPay: Domain-Adaptive LoRA Fine-Tuning of Whisper for Low-Resource Nepali Financial Speech Recognition

Read the original on arXiv Computation and Language →

SpeakPay is a voice‑first digital wallet designed to make mobile payment apps in Nepal accessible to visually impaired users. The paper introduces NepFinSpeech‑403, a 403‑utterance Nepali financial voice command dataset, and demonstrates that fine‑tuning Whisper large‑v2 with LoRA reduces the Word Error Rate from 129.95% to 42.58% and improves Devanagari numeral recognition from 0.0% to 73.9%. Domain adaptation also boosts the Transaction Success Rate from 1.67% to 33.33%, with as few as 100 domain‑specific utterances halving the zero‑shot WER.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Aug 27

Fine-Tuning Whisper for Automatic Speech Recognition in Baniwa: A Preliminary Study

The study fine‑tunes the Whisper Small model for Automatic Speech Recognition (ASR) in Baniwa, an indigenous Arawakan language. Using a 0.54‑hour corpus of 1,373 manually transcribed recordings, the fine‑tuned model achieved a Word Error Rate of 37.5% and a Character Error Rate of 7.45%. These results provide an initial baseline for Baniwa ASR and suggest that multilingual foundation models can be adapted to extremely low‑resource languages.

By Leonardo Duart, Tiago Fonseca, Thiago Chac\'on
arXiv Computation and Language
Aug 24

Building and Evaluating a Synthetic Bengali Speech Resource for Telecom Customer Care

The paper introduces a synthetic Bengali speech dataset tailored for telecom customer‑care applications, comprising 10,000 audio‑text pairs (≈26.82 hours) with predefined train, validation, and test splits. The data were generated using OmniVoice voice‑cloning, and include both original and normalized transcripts for ASR/STT use. Automatic intelligibility evaluation with a fine‑tuned Whisper model shows an average WER of 2.54% and CER of 0.59%, indicating strong text‑audio consistency, while the authors note limitations of synthetic speech and STT‑based evaluation.

By Kawshik Kumar Paul, Md. Nafiul Alam Fuji