The paper investigates how large language models (LLMs) can be used to generate synthetic data for low‑resource machine translation, focusing on Romansh, which has six distinct varieties. It finds that LLMs are better at translating from Romansh to a high‑resource language (German) than the reverse, creating an asymmetry that makes the direction of data augmentation critical. By generating synthetic translations into German rather than Romansh, the authors surpass a Gemini 3 Pro baseline on German‑Romansh translation, achieving a +23 BLEU improvement in the lowest‑resource variety and producing fluent translations in each Romansh variety according to human evaluation.
By Jannis Vamvas, Ignacio P\'erez Prat, Angela Heldstab, Dominic P. Fischer, Sina Ahmadi, Rico Sennrich
The paper examines two test‑time scaling methods for large language models in machine translation: sequential sampling, where later attempts build on earlier ones, and parallel sampling, such as independent i.i.d. sampling with reranking. Sequential sampling shows a higher performance ceiling, offering a more diverse and effective set of translations, especially with limited sampling budgets. Human analysis reveals that while sequential sampling improves fluency and naturalness, it can reduce accuracy when the inference budget is large, and the authors attribute this effect to the model’s access to a larger target‑side context.
By Di Wu, Sergey Troshin, Christof Monz, Antske Fokkens, Vlad Niculae
arXiv:2608. 09766v1 Announce Type: cross Abstract: Multilingual translation benchmarks are typically sourced in English and translated into other languages, treating language pairs as the unit of evaluation---a design that is prone to contamination over time and overlooks locale and cultural considerations.
By Pinzhen Chen, Koel Dutta Chowdhury, Xiaoya Xu, David Tan, Doreen Osmelak, Ona de Gibert, Ariun-Erdene Tumurchuluun, Ashok Urlana, Fedor Sizov, Hale Sirin, Jesujoba Alabi, Karrar Talib Abed, Mateusz Klimaszewski, Nikolay Bogoychev, Niyati Bafna, Patricia Schmidtova, Preksha Manjunath Shanbhag, Sherrie Shen, Vilem Zouhar, Vivek Iyer, Yasser Hamidullah, Yusser Al Ghussin, Zheng Zhao
EnSiTa is a trilingual multi‑domain parallel dataset and benchmark for English, Sinhala, and Tamil. It contains human post‑edited training data across seven domains and professionally translated test sets for those domains plus an additional one, all produced through a multi‑year, rigorously quality‑controlled process. The authors use EnSiTa to conduct a comprehensive study of domain‑specific machine translation across six language directions, comparing from‑scratch Transformers, pre‑trained models, and decoder‑only LLMs under various training‑data sizes, model scales, and domain settings.
By Surangika Ranathunga, Nisansa de Silva, Aloka Fernando, Kavindu Warnakulasuriya, Isuru Wijesiri, Menan Velayuthan, Charitha Rathnayaka, Thivaharan Varatharajan, Sajeevi Silva, Piumi Kandanaarachchi, Uthayasanker Thayasivam
EuroAlpaca presents a task‑preserving localisation pipeline that translates English instruction‑tuning data into 50 European languages while maintaining task‑critical constraints. The method uses field‑wise machine translation or reconstructs task‑equivalent target‑language instances, followed by validation of coherence and consistency. Experiments show that EuroAlpaca improves instruction‑following accuracy by 12.9% over a baseline and outperforms direct translation on ROUGE‑L and F‑BERT metrics.
By Aleix Sant, Jordi Luque, Carlos Escolano
Fine‑tuning large language models on parallel data can improve translation quality but also causes catastrophic forgetting of general capabilities. The study evaluates several forgetting‑mitigation methods—anchored to auxiliary data, model outputs, and base model parameters—using Llama 3.2 1B Instruct and Llama 3.1 8B Instruct on Arabic‑English and Spanish‑English translation tasks. Elastic Weight Consolidation best preserves general benchmark performance, yet only data mixing with control‑task examples maintains instruction‑following abilities such as formality and grammatical gender control, though these gains do not generalize to unseen prompts.
By Niklas Scholz, David Thulke, Abdallah Nasir, Will Allred, Evgeny Matusov, Hermann Ney
arXiv:2608. 08283v1 Announce Type: cross Abstract: Although large language models can translate some historical languages surprisingly well, their usefulness in digital humanities workflows is limited by the lack of reliable evaluation.
By Osvaldo Quinjica, Eric Bennett, Xinchen Yang, Andrew Schonebaum, Marine Carpuat
arXiv:2609.18720v1 Announce Type: new
Abstract: Learned quality estimation (QE) models such as COMETKiwi are widespread and work well for general machine translation evaluation. However, they are kno...
By Kathy H\"ammerl, Gabriel Bretschner, Joern Wuebker
The paper introduces ReMova, a pipeline for cleaning Belarusian data and fine‑tuning large language models (LLMs) for English‑to‑Belarusian translation. It uses a correction tool to handle the two orthographies of Belarusian, remove noise, filter out interference from other languages, and correct common misspellings found online. Ablation experiments on unfiltered data show that filtering benefits all fine‑tuned models, with LLM‑based models gaining about twice as much as a dedicated encoder‑decoder MT system, highlighting data quality as a key bottleneck for Belarusian MT.
By Mikita Pilinka, Aliaksandr Kliuje\u{u}, David Samuel, Yves Scherrer
arXiv:2609.16340v1 Announce Type: cross
Abstract: Machine translation systems are periodically upgraded to stronger models, but the available preference signal is human post-edits of an older system'...
By Rohit Dhaipule, Sukhdeep Singh Kharbanda, Prasanth Bathala, Pradyumna Lanka, Anubhav Shrimal
arXiv:2609.06634v1 Announce Type: cross
Abstract: LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they...
By Lifeng Han, Jiahui Liang, Anna Latusek, Karim El Haff, Amal Haddad Haddad, Josua H\"ofgen, Kilian Evang, Min Ma, Maryia Zhyrko
arXiv:2608.03446v2 Announce Type: replace
Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the giv...
By Adnan Al Ali, Kathy H\"ammerl, Jind\v{r}ich Libovick\'y, Alexander Fraser