Hugging Face Trending Papers

CoPiT: Cognitive Pivot Translation for Digraphic Low-Resource Mongolian in the Traditional Script

Low-resource languages remain challenging for machine translation, and Mongolian is a representative case. As a digraphic language, Mongolian is written in both Cyrillic and Traditional scripts, which exhibit a severe imbalance in data availability.

arXiv Computation and Language
Sep 11

A Factorial Study of Synthetic Data Generation for Low-Resource Machine Translation using Grammar Books

The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.

By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
arXiv Computation and Language
Sep 25

EnSiTa - A Trilingual Multi-Domain Parallel Dataset and Benchmark for Domain-Specific Machine Translation

EnSiTa is a trilingual multi‑domain parallel dataset and benchmark for English, Sinhala, and Tamil. It contains human post‑edited training data across seven domains and professionally translated test sets for those domains plus an additional one, all produced through a multi‑year, rigorously quality‑controlled process. The authors use EnSiTa to conduct a comprehensive study of domain‑specific machine translation across six language directions, comparing from‑scratch Transformers, pre‑trained models, and decoder‑only LLMs under various training‑data sizes, model scales, and domain settings.

By Surangika Ranathunga, Nisansa de Silva, Aloka Fernando, Kavindu Warnakulasuriya, Isuru Wijesiri, Menan Velayuthan, Charitha Rathnayaka, Thivaharan Varatharajan, Sajeevi Silva, Piumi Kandanaarachchi, Uthayasanker Thayasivam
arXiv Computation and Language
Sep 25

Benchmarking Arabic--Russian Machine Translation: A Comparison of Fine-tuned NMT and Few-shot LLMs under Rich Morphology and Low Lexical Overlap

The study evaluates Arabic–Russian machine translation by comparing seven fine‑tuned neural machine translation (NMT) models with four few‑shot large language models (LLMs) on a new 15.47 million‑pair corpus split into 20k/5k/5k. Fine‑tuned NLLB‑1.3B achieves the best performance (BLEU 16.3, COMET 0.738), while the best few‑shot LLM, Aya‑Expanse 8B, scores only BLEU 1.7 on 500 sentences. Error analysis shows that low lexical overlap between Arabic and Russian is the main source of failures, and statistical tests confirm significant performance gaps between most models.

By Mullosharaf K. Arabov
arXiv AI
Sep 4

Translation as a Decision Space: A Multi-Agent Perspective on Low-Resource Dialect Generation

The paper proposes treating translation as a structured decision space explored by multiple autonomous agents, rather than producing a single output. Using Turkish–Syrian Arabic dialogue, three agents—zero‑shot, dialect‑stabilized, and pivot translation—are compared on 5,000 sentences, with stabilization nearly doubling dialect marker usage and reducing structural instability. The study introduces an interpretability framework that quantifies decision flexibility through dialect marker frequency, lexical proximity, and structural variance.

By Hasan Alkhder, Mohammad Abboush, Igor Tchappi, Ahmet Zengin, Amro Najjar
arXiv Machine Learning
Aug 11

Embedding Initialization for Unseen Low-resource Languages in Multilingual NMT: A Case Study on Limbum-English Translation

arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.

By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv Computation and Language
Sep 15

North Small Translate: Advanced Cost-Effective Translation (Cohere CAT+)

arXiv:2609.13916v1 Announce Type: new Abstract: We present North Small Translate, an open-weight, LLM-based machine translation (MT) model with instruction-following capabilities built on the same fo...

By Tom Kocmi, Alexandre B\'erard, Phil Blunsom, Samuel Cahyawijaya, Shaun Cassini, Nicholas Frosst, Ona de Gibert, Aidan Gomez, Nithya Govindarajan, Shun Kiyono, Olivia Lasche, Lawrence Rogers, Kelly Marchisio, Nikita Moghe, Yash More, Camila Moran-Hidalgo, Yiyang Nan, Michael Sachs, Trisha Starostina, Daan van Stigt, Spencer Rarrick, Sebastian Vincent, Ivan Zhang
arXiv AI
Sep 4

One Model to Translate Them All? A Journey to Mount Doom for Multilingual Model Merging

The paper investigates weight‑space merging of independently fine‑tuned multilingual machine translation models. Experiments show that merging is more successful when models share a target language, yet it still cannot match the peak performance of language‑specific checkpoints. When target languages differ, performance drops sharply, and analysis reveals that overlapping neuron activation and incompatible upper‑layer geometries cause these failures.

By Baban Gain, Trilok Nath Singh, Asif Ekbal