arXiv Computation and Language By Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

Read the original on arXiv Computation and Language →

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a family of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can remove up to 96% of training tokens without harming quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this mixture outperform much larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

Hugging Face Trending Papers
Aug 19

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a suite of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can discard up to 96% of training tokens while preserving quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this curated mixture outperform larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

arXiv Machine Learning
Aug 11

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.

By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli