arXiv Machine Learning

Sample-Size Scaling of the African Languages NLI Evaluation

arXiv:2606. 03219v1 Announce Type: cross Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance.

arXiv AI
Jul 7

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

arXiv:2607. 04814v1 Announce Type: cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale.

By Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba
arXiv Computation and Language
Aug 28

AfriSwitch: A Benchmark for In-the-Wild African Code-Switched Speech Recognition

AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.

By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
Hugging Face Trending Papers
Jun 2

From Script to Semantics: Prompting Strategies for African NLI

Large language models (LLMs) are increasingly evaluated in multilingual settings, yet their inference behavior in low-resource African languages remains underexplored especially under pure prompting without fine-tuning. We present a systematic study of prompting strategies for Natural Language Inference (NLI) in Swahili, Yoruba, and Hausa using the AfriXNLI benchmark.

arXiv Computation and Language
Aug 27

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a family of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can remove up to 96% of training tokens without harming quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this mixture outperform much larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

By Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir
arXiv Machine Learning
Aug 11

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.

By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv Computation and Language
Aug 31

Ladders in Chaos: When, How, (and Perhaps Why) Does Test-Time Scaling Improve LLM Machine Translation

The paper examines two test‑time scaling methods for large language models in machine translation: sequential sampling, where later attempts build on earlier ones, and parallel sampling, such as independent i.i.d. sampling with reranking. Sequential sampling shows a higher performance ceiling, offering a more diverse and effective set of translations, especially with limited sampling budgets. Human analysis reveals that while sequential sampling improves fluency and naturalness, it can reduce accuracy when the inference budget is large, and the authors attribute this effect to the model’s access to a larger target‑side context.

By Di Wu, Sergey Troshin, Christof Monz, Antske Fokkens, Vlad Niculae
arXiv Computation and Language
Aug 27

Apples to Apples? Towards Comparable Crosslingual Language Model Evaluation

The paper investigates how to fairly compare language models across languages, noting that current evaluation methods vary widely and lack empirical validation. By training controlled monolingual models on parallel data and testing multilingual LLMs, the authors find that many normalized metrics suffer from biases due to tokenization, encoding, and orthographic differences. Instead, they recommend using sentence‑level negative log‑likelihood over semantically equivalent sequences for more reliable cross‑lingual comparisons.

By Xiulin Yang, Ethan Gotlieb Wilcox, Catherine Arnett
arXiv AI
Sep 2

The Curse of Multilinguality in Lexical Normalization

The paper investigates how many languages should be jointly trained in a single lexical normalization model. Using a fixed-capacity character-level model across twelve languages, it finds that accuracy peaks when a language is trained with only a few others—typically one to four—and then declines sharply as more languages are added, dropping about forty percent. A control experiment keeping total training data constant shows the decline is due to competition for model capacity rather than data scarcity, and no reliable typological rule predicts the optimal number of co‑training languages.

By Saman Rahbar