arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.
By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv:2607. 20241v1 Announce Type: cross Abstract: Culturally loaded translation poses unique challenges for machine translation (MT), as meanings are deeply embedded in socio-cultural contexts beyond surface linguistic forms.
By Yiming Wang, Jiayuan Di
The paper investigates how large language models (LLMs) can be used to generate synthetic data for low‑resource machine translation, focusing on Romansh, which has six distinct varieties. It finds that LLMs are better at translating from Romansh to a high‑resource language (German) than the reverse, creating an asymmetry that makes the direction of data augmentation critical. By generating synthetic translations into German rather than Romansh, the authors surpass a Gemini 3 Pro baseline on German‑Romansh translation, achieving a +23 BLEU improvement in the lowest‑resource variety and producing fluent translations in each Romansh variety according to human evaluation.
By Jannis Vamvas, Ignacio P\'erez Prat, Angela Heldstab, Dominic P. Fischer, Sina Ahmadi, Rico Sennrich
TransClean introduces a benchmark for identifying and removing translation noise—unwanted text such as language labels, explanations, or bilingual repetitions—from large language model (LLM) outputs. The authors analyzed 790,000 translations from 12 LLMs across 22 language pairs, cataloguing 12 common noise patterns and creating 9,900 noisy‑clean pairs (8,800 synthetic, 1,100 authentic). They evaluated two extraction methods—a span‑based approach using quality estimation models and an LLM‑prompted method—demonstrating the first systematic framework to assess and improve translation cleanliness.
By Shenbin Qian, Yves Scherrer
The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.
By Baban Gain, Trilok Nath Singh, Asif Ekbal
arXiv:2508. 05502v2 Announce Type: replace-cross Abstract: Multimodal Large Language Models (MLLMs) perform strongly in high-resource languages, yet often produce fluent but culturally "thin" descriptions in low-resource settings.
By Yufei Gao, Jiaying Fei, Nuo Chen, Ruirui Chen, Guohang Yan, Yunshi Lan, Botian Shi
arXiv:2609.18720v1 Announce Type: new
Abstract: Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are kno...
By Kathy H\"ammerl, Gabriel Bretschner, Joern Wuebker
arXiv:2605.28190v2 Announce Type: replace
Abstract: Embedding benchmarks like MTEB report a single score per model, implicitly treating robustness as a static, scalar property. We argue that embeddin...
By Manuel Frank, Haithem Afli
Dialectal variation remains a major challenge for multilingual language models. Perturbation-based continued pre-training (CPT) has emerged as a promising approach to improving robustness, yet existing work largely evaluates individual perturbation strategies in isolation and provides limited insight into why they work.
arXiv:2606. 28843v1 Announce Type: cross Abstract: Fine-tuning a large language model is a ubiquitous method for enhancing its capability on a specific downstream task.
By Will Hawkins, Kaivalya Rawal, Jonathan Rystr{\o}m, Stratis Tsirtsis, Zihao Fu, Greta Warren, Ryan Brown, Eoin Delaney, Sandra Wachter, Brent Mittelstadt, Chris Russell
The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.
By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
The study evaluates Arabic–Russian machine translation by comparing seven fine‑tuned neural machine translation (NMT) models with four few‑shot large language models (LLMs) on a new 15.47 million‑pair corpus split into 20k/5k/5k. Fine‑tuned NLLB‑1.3B achieves the best performance (BLEU 16.3, COMET 0.738), while the best few‑shot LLM, Aya‑Expanse 8B, scores only BLEU 1.7 on 500 sentences. Error analysis shows that low lexical overlap between Arabic and Russian is the main source of failures, and statistical tests confirm significant performance gaps between most models.
By Mullosharaf K. Arabov