AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.
By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
arXiv:2606.06037v3 Announce Type: replace-cross
Abstract: Large audio language models (LALMs) are increasingly deployed in real-world applications, yet their safety alignment is still primarily evalu...
By Virginia Ceccatelli, Yejin Jeon, David Ifeoluwa Adelani
arXiv:2609.09554v1 Announce Type: new
Abstract: We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. L...
By Shivam Singh, Aditya Yadavalli, Catherine Arnett, Alex Warstadt
We introduce BuzzASR, a collection of language-specialized fine-tuned Whisper models adapted for automatic speech recognition (ASR) in 102 languages. Large end-to-end Transformer-based ASR models such...
The paper addresses the challenge of identifying code‑switched utterances in language identification systems. It revisits the MaskLID approach, highlighting its overreliance on word‑level language association scores, and reformulates its optimization as an Integer Linear Program to incorporate clear, interpretable constraints. These enhancements significantly improve performance across ten diverse languages on code‑switching benchmarks, with the authors releasing code and data for reproducibility.
By Joanna Rado{\l}a, Josep Maria Crego, Fran\c{c}ois Yvon
We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2. 0 self-supervised speech encoder.