arXiv Computation and Language By Sewade Ogun

Sometin Beta Pass Notin: Improving Multilingual ASR for Nigerian Languages via Knowledge Distillation

Read the original on arXiv Computation and Language →

The paper presents Sometin Beta Pass Notin (SBPN), a multilingual ASR framework for Nigerian languages that uses a two‑stage knowledge‑distillation approach. First, student‑teacher distillation from monolingual models is conditioned on language‑specific N‑gram language models; second, iterative self‑improvement with pseudo‑labelled data further refines accuracy. The method reduces relative WER by 29% over monolingual baselines and outperforms state‑of‑the‑art multilingual models on Common Voice and FLEURS benchmarks for Yoruba, Hausa, Igbo, Nigerian Pidgin, and Nigerian English.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Jul 7

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

arXiv:2607. 04814v1 Announce Type: cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale.

By Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba
arXiv Computation and Language
Sep 22

Cross-Dialect NER for Bangla Regional Dialects Using Leave-One-Dialect-Out Cross-Validation and Explainable AI

The paper introduces a cross-dialect Named Entity Recognition (NER) framework for Bangla, leveraging the ANCHOLIK-NER dataset that covers five major regional dialects. Using a Leave-One-Dialect-Out Cross-Validation strategy, eight transformer-based models were evaluated, with Multilingual-E5 Large achieving the best performance (F1 up to 97.26% on Mymensingh, 82.38% on Chattogram). Local Interpretable Model-agnostic Explanations (LIME) revealed that the models rely mainly on the surface form of entity words rather than surrounding context, suggesting a direction for future improvement.

By Shamim Rahim Refat, Faika Fairuj Preotee, Shuvashis Sarker, Shifat Islam, Bidyarthi Paul, Mohammad Ashraful Hoque