arXiv Computation and Language

ReMova: Fine-tuning LLMs for English to Belarusian translation

The paper introduces ReMova, a pipeline for cleaning Belarusian data and fine‑tuning large language models (LLMs) for English‑to‑Belarusian translation. It uses a correction tool to handle the two orthographies of Belarusian, remove noise, filter out interference from other languages, and correct common misspellings found online. Ablation experiments on unfiltered data show that filtering benefits all fine‑tuned models, with LLM‑based models gaining about twice as much as a dedicated encoder‑decoder MT system, highlighting data quality as a key bottleneck for Belarusian MT.

arXiv Computation and Language
Aug 27

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a family of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can remove up to 96% of training tokens without harming quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this mixture outperform much larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

By Milan Gritta, Patrik Lambert, Jihye Back, Amril Nazir
Hugging Face Trending Papers
Aug 19

TranslatePsy-AfriSLM: High-Quality Data Scaling For Low-Resource Machine Translation

TranslatePsy-AfriSLM is an open‑source machine‑translation resource set for 19 Sub‑Saharan African languages, comprising curated parallel data, African‑specialized synthetic data, and a suite of fine‑tuned small language models (SLMs). The authors demonstrate that a unified quality‑estimation filtering can discard up to 96% of training tokens while preserving quality, and that filtered synthetic data dominates the quality‑efficiency Pareto frontier. Models trained on this curated mixture outperform larger systems such as TranslateGemma‑27B and Qwen3.5‑122B‑A10B, achieving superior performance with as few as 0.8 B parameters.

arXiv Computation and Language
Sep 25

Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap

The study evaluates Arabic–Russian machine translation by comparing seven fine‑tuned neural machine translation (NMT) models with four few‑shot large language models (LLMs) on a new 15.47 million‑pair corpus split into 20k/5k/5k. Fine‑tuned NLLB‑1.3B achieves the best performance (BLEU 16.3, COMET 0.738), while the best few‑shot LLM, Aya‑Expanse 8B, scores only BLEU 1.7 on 500 sentences. Error analysis shows that low lexical overlap between Arabic and Russian is the main source of failures, and statistical tests confirm significant performance gaps between most models.

By Mullosharaf K. Arabov
arXiv Machine Learning
Sep 24

Fine-Tuning LLMs for Translation: General Forgetting Mitigation Does Not Preserve MT-Specific Instruction Following

Fine‑tuning large language models on parallel data can improve translation quality but also causes catastrophic forgetting of general capabilities. The study evaluates several forgetting‑mitigation methods—anchored to auxiliary data, model outputs, and base model parameters—using Llama 3.2 1B Instruct and Llama 3.1 8B Instruct on Arabic‑English and Spanish‑English translation tasks. Elastic Weight Consolidation best preserves general benchmark performance, yet only data mixing with control‑task examples maintains instruction‑following abilities such as formality and grammatical gender control, though these gains do not generalize to unseen prompts.

By Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney
arXiv Computation and Language
Sep 11

TransClean: A Benchmark for Detecting and Extracting Clean Translations from Large Language Model Outputs

TransClean introduces a benchmark for identifying and removing translation noise—unwanted text such as language labels, explanations, or bilingual repetitions—from large language model (LLM) outputs. The authors analyzed 790,000 translations from 12 LLMs across 22 language pairs, cataloguing 12 common noise patterns and creating 9,900 noisy‑clean pairs (8,800 synthetic, 1,100 authentic). They evaluated two extraction methods—a span‑based approach using quality estimation models and an LLM‑prompted method—demonstrating the first systematic framework to assess and improve translation cleanliness.

By Shenbin Qian, Yves Scherrer
arXiv Computation and Language
Sep 21

Dynamic Lagging using Stable-Prefix Training for Simultaneous Translation

The paper introduces a training strategy for cascaded simultaneous speech translation that allows the system to dynamically decide how much of the source prefix to translate. By fine‑tuning a large language model (Qwen3‑8B) on stable prefixes—pairs of source prefixes and the longest shared translation with the full sentence—the authors enable contextual read‑write decisions beyond fixed wait‑k or target‑suffix deletion. Experiments on English‑to‑German, Japanese, and Chinese demonstrate that stable prefixes improve the quality‑latency tradeoff across various test sets.

By Hieu Hoang, Amittai Axelrod, Matt Post
arXiv Computation and Language
4d ago

Index-Translate: A Multilingual Translation Model Family -- Text, Speech, Controlled Dubbing, and Long-Document Translation

arXiv:2609.40181v1 Announce Type: new Abstract: We introduce Index-Translate, a multilingual translation model family that combines a shared multilingual foundation with specialized training for gene...

By Tianjiao Li, Mengran Yu, Chenyu Shi, Lusheng Zhang, Qisi Chen, Yanshan Zhou, Ji Qi, Jingying Liu, Yuang Feng, Ziang Cui, Tianxing Yan
arXiv AI
Sep 25

TTLab at AlexandriaX-2026: A Fine-Tuned Surface Tagger for Arabic Machine-Translation Error-Span Detection and Classification

TTLab submitted a system for the AlexandriaX-2026 Subtask 3 on Arabic machine‑translation error‑span detection and classification. The approach treats the task as token‑level classification over surface forms, using focal loss with class weighting and dialect‑specific decoding thresholds to address label imbalance. MARBERTv2, among six Arabic pre‑trained encoders, achieved the best performance, ranking third overall, though classification of rare error types remains difficult, indicating a need for data augmentation.

By Ali Abusaleh, Bhuvanesh Verma, Alexander Mehler