AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.
By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
arXiv:2609.15758v1 Announce Type: new
Abstract: Extending large-scale multilingual automatic speech recognition (ASR) models to low-resource languages remains challenging. Model performance is skewed...
By Thai Thi Thanh Thao Dang, Mengjie Qian, Kate Knill
The study evaluates how different input representations—orthographic text, IPA transcription, and romanization—affect cross‑lingual transfer in autoregressive multilingual language models. Across three model sizes and eight languages grouped into typologically motivated pairs, romanized pretraining consistently outperforms native orthography and IPA, especially as model scale increases. Fine‑tuning a text‑pretrained model on romanized data can harm performance on languages already covered by the base model, suggesting romanization should be integrated at pretraining rather than applied later.
By Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar
arXiv:2606. 03219v1 Announce Type: cross Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance.
By Anuj Tiwari, Oluwapelumi Ogunremu, Terry Oko-odion, Jesujuwon Egbewale, Hannah Nwokocha
arXiv:2609.37543v1 Announce Type: new
Abstract: Cross-lingual zero-shot transfer and multilingual fine-tuning are promising approaches for NLP tasks such as Named Entity Recognition (NER) in low-reso...
By Prosper Arineitwe Asiimwe, Francois Meyer, Jan Buys
arXiv:2510. 15551v2 Announce Type: replace-cross Abstract: Any piece of knowledge is usually expressed in one or a handful of natural languages on the web or in any large corpus.
By Vihari Piratla, Purvam Jain, Darshan Singh, Trevor Cohn, Preethi Jyothi, Partha Talukdar