arXiv:2606.30943v2 Announce Type: replace
Abstract: Russian and Arabic are among the major languages of scientific communication. Language barriers impede the exchange of research results between the...
By Mullosharaf K. Arabov
The study evaluates Arabic–Russian machine translation by comparing seven fine‑tuned neural machine translation (NMT) models with four few‑shot large language models (LLMs) on a new 15.47 million‑pair corpus split into 20k/5k/5k. Fine‑tuned NLLB‑1.3B achieves the best performance (BLEU 16.3, COMET 0.738), while the best few‑shot LLM, Aya‑Expanse 8B, scores only BLEU 1.7 on 500 sentences. Error analysis shows that low lexical overlap between Arabic and Russian is the main source of failures, and statistical tests confirm significant performance gaps between most models.
By Mullosharaf K. Arabov
arXiv:2609.17539v1 Announce Type: new
Abstract: We present MudawanSn, a gold-standard resource of 1,271 sentence-aligned pairs manually translated from Wolof into Modern Standard Arabic (MSA). The so...
By Mouhamed Mbaye, Thierno Diop
arXiv:2609.06634v1 Announce Type: cross
Abstract: LLMs' performance on machine translation (MT) tasks is often dependent on the data availability in the specific domains and language pairs that they...
By Lifeng Han, Jiahui Liang, Anna Latusek, Karim El Haff, Amal Haddad Haddad, Josua H\"ofgen, Kilian Evang, Min Ma, Maryia Zhyrko
BEAR-Bench is a bilingual benchmark for multimodal large language models, featuring 1,000 human‑annotated questions derived from text‑rich business and scientific documents in English and Russian. It evaluates 16 MLLMs, including Gemini 3.1 Pro and Qwen3.5‑397B, revealing significant performance gaps even for the strongest systems. The benchmark also serves to compare hallucination‑detection methods by analyzing model failures on these complex documents.
By Liubov Chubarova, Alexandra Kuleshova, Daniil Volkov, Kirill Sultanov, Alexey Zaytsev
arXiv:2608.12018v2 Announce Type: replace
Abstract: Neural Machine Translation (NMT) and Large Language Models (LLMs) excel at cross-lingual tasks but often fail to capture intra-lingual morphologica...
By Rakib Ullah, Md. Ruhul Islam, Tanbir Ahmed, Nayan Kumar Nath
arXiv:2011.03783v3 Announce Type: replace-cross
Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with an...
By Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
arXiv:2607.00890v2 Announce Type: replace
Abstract: Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synth...
By Maximilian Idahl, J\"org Tiedemann, Sampo Pyysalo, David Salinas, Tomasz Galica, Shenbin Qian, Tudor Nicolae Mateiu, Zihao Li, Anna Lokrantz, Fedor Vitiugin, Andr\'e F. T. Martins, Jenna Kanerva, Filip Ginter, Matthias Lindemann, Tim Isbister, Birger Moell, Jonas Lindh, Jan Haji\v{c}, Jenia Jitsev, Andrey Kutuzov, Stephan Oepen, Gema Ram\'irez-S\'anchez
Open web-scale pre-training corpora remain concentrated in English, limiting multilingual LLM development. We introduce MultiSynt/MT, an open synthetic parallel corpus with approximately 4.
arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.
By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv:2509. 07829v4 Announce Type: replace-cross Abstract: Literary translation has recently gained attention as a distinct and complex task in machine translation research, yet translation by small open models remains an open problem, particularly for low-resource languages such as Romanian.
By Mihai Nadas, Laura Diosan, Andreea Tomescu, Andrei Piscoran