arXiv:2607.00890v2 Announce Type: replace
Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synth...
By Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Haji\v{c}, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ram\'irez-S\'anchez
arXiv:2608.12018v2 Announce Type: replace
Abstract: Neural Machine Translation (NMT) and Large Language Models (LLMs) excel at cross-lingual tasks but often fail to capture intra-lingual morphologica...
By Rakib Ullah, Md. Ruhul Islam, Tanbir Ahmed, Nayan Kumar Nath
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.
arXiv:2609.06634v1 Announce Type: cross
Abstract: LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they...
By Lifeng Han, Jiahui Liang, Anna Latusek, Karim El Haff, Amal Haddad Haddad, Josua H\"ofgen, Kilian Evang, Min Ma, Maryia Zhyrko
arXiv:2609.13615v1 Announce Type: new
Abstract: For our submission to the WMT26 Creole Language Translation Shared Task, we focus on machine translation (MT) models for Pacific creoles: Tok Pisin, Bi...
By Rapha\"el Merx, Nick Thieberger, Ekaterina Vylomova
arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.
By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv:2608.03446v2 Announce Type: replace
Abstract: Multilingual large language models (LLMs) have been shown to perform better on non-English classification tasks when the representations of the giv...
By Adnan Al Ali, Kathy H\"ammerl, Jind\v{r}ich Libovick\'y, Alexander Fraser
The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.
By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
COILD is an Indic‑centric parallel corpus that contains over 1.16 million human‑translated and verified sentence pairs across 20 Indian language pairs from four language families. The corpus is sourced from original Indian language materials in eight domains, and a 2,000‑sentence domain‑centric benchmark is provided for consistent multilingual evaluation. Experiments with IndicTrans2‑Distilled and NLLB‑200 show consistent improvements in automatic metrics and human judgments, underscoring the value of high‑quality Indic‑centric data.
By Kshetrimayum Boynao Singh, Nitin Kumar Mishra, Palash Pratim Dutta, Atai Waris Khan, Aparna Kaushik, Avinash Kumar, Deeksha, Deepak Kumar, Saroj Kumar Jha, Saloka Sengupta, Anansa Roy, Umalatha Kannoth, Saifulla Samar, Meena Sharma, Manpreet Kaur, Jyoti Sharma, Ashwini Vaidya, Muralikrishna SN, Md Shad Akhtar, Poonam Bansal, Amita Dev, Sanasam Ranbir Singh, Samit Bhattacharya, Tanmoy Chakraborty, Asif Ekbal
arXiv:2602.14488v3 Announce Type: replace-cross
Abstract: IR in low-resource languages remains limited by the scarcity of high-quality, task-specific annotated datasets. Manual annotation is expensiv...
By Md. Najib Hasan, Mst. Jannatun Ferdous Rain, Fyad Mohammed, Nazmul Siddique
arXiv:2606. 13647v1 Announce Type: cross Abstract: We introduce SkMTEB, the first comprehensive MTEB-style text embedding benchmark for Slovak, a low-resource West Slavic language, comprising 31 datasets across 7 task types -- nearly 4$\times$ the depth of existing multilingual benchmark coverage for Slovak.
By Marek \v{S}uppa, Andrej Ridzik, Daniel Hl\'adek, Nat\'alia K\v{n}a\v{z}ekov\'a, Vikt\'oria Ondrejov\'a
arXiv:2509.17930v3 Announce Type: replace-cross
Abstract: Multilingual translation suffers from computational redundancy, especially when translating into multiple languages simultaneously. In additi...
By Yiwen Guan, Jacob Whitehill