arXiv AI By Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

Read the original on arXiv AI →

arXiv:2607. 04814v1 Announce Type: cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 28

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.

By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
arXiv Computation and Language
Aug 27

One Form to Transfer Them All: Pretraining Multilingual Language Models Beyond Native Orthography

The study evaluates how different input representations—orthographic text, IPA transcription, and romanization—affect cross‑lingual transfer in autoregressive multilingual language models. Across three model sizes and eight languages grouped into typologically motivated pairs, romanized pretraining consistently outperforms native orthography and IPA, especially as model scale increases. Fine‑tuning a text‑pretrained model on romanized data can harm performance on languages already covered by the base model, suggesting romanization should be integrated at pretraining rather than applied later.

By Muge Zhang, Aaron Jencks, Krishna Badikela, Yulia Tsvetkov, Sachin Kumar