Hugging Face Trending Papers

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a suite of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can discard up to 96% of training tokens while preserving quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this curated mixture outperform larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

arXiv Computation and Language
Aug 27

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a family of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can remove up to 96% of training tokens without harming quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this mixture outperform much larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

By Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir
arXiv Machine Learning
Aug 11

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.

By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv Computation and Language
Aug 25

LuxIT: A Luxembourgish Instruction Tuning Dataset from Monolingual Seed Data

LuxIT is a monolingual instruction‑tuning dataset for Luxembourgish, created by synthesizing instruction‑answer pairs from native texts using the DeepSeek‑R1‑0528 model and a quality‑assurance LLM‑as‑judge process. The resulting 227,507 high‑quality pairs were used to fine‑tune 14 LLMs (≤15 B parameters), yielding an average accuracy increase of +5.37 percentage points on standardized Luxembourgish proficiency exams and improvements in macro‑averaged F1 on nine of the fourteen downstream NLP tasks. These findings demonstrate that synthetic monolingual data can effectively enhance LLM performance in low‑resource languages and reveal the complex relationship between exam performance and practical NLP gains.

By Julian Valline, Cedric Lothritz, Siwen Guo, Jordi Cabot
arXiv AI
Jul 13

A Sovereign, Open-Source Foundation Model for German and English

arXiv:2607. 09424v1 Announce Type: cross Abstract: We present Soofi S 30B-A3B, a sovereign, open-source Mixture-of-Experts (MoE) hybrid Mamba Transformer foundation model for German and English.

By The Soofi-Team, :, Benedikt Droste, David Fitzek, Ruben H\"arle, Lukas Helff, Maximilian Idahl, Alex Jude, Abbas Goher Khan, Maurice Kraus, Timm Ruland, Richard Rutmann, Sebastian Sztwiertnia, Markus Frey, Daniil Gurgurov, Jan Pfister, Tom R\"ohr, Sebastian von Rohrscheidt, J\"org Bienert, Nicolas Flores-Herr, Simon Gottschalk, Andreas Hotho, Kristian Kersting, Joachim K\"ohler, Alexander L\"oser, Wolfgang Nejdl, Simon Ostermann, Jan Plogsties, Patrick Putzky, Mehdi Ali, Michael Fromm, Max L\"ubbering
arXiv AI
Jul 7

Evaluating the Effect of Linguistic Relatedness on Cross-Lingual Transfer in Large Multilingual Automatic Speech Recognition

arXiv:2607. 04814v1 Announce Type: cross Abstract: Extending automatic speech recognition (ASR) to low-resource African languages is constrained by the prohibitive demands of data collection at scale.

By Andrei Florian, Cynthia Jayne Amol, Hope Kerubo Ombaba, Xiaoyu Cui, Boniface Mwau, Biatus Maina Kamau, Lilian Diana Awuor Wanzare, Christiane Fellbaum, Happy Buzaaba