arXiv:2609.18529v1 Announce Type: new
Abstract: UNESCO considers the Assyrian (Syriac) language an endangered language. Although Assyrians speak the language worldwide, the speaking population is unc...
By Hadiana Sliwa, Hossein Hassani
arXiv:2606. 25365v2 Announce Type: replace-cross Abstract: We present a study on low-resource machine translation for the Tangkhul-English (nmf-en) language pair.
By Chormi Zimik Vashai, Agniva Maiti
arXiv:2607. 23319v1 Announce Type: cross Abstract: Standard subword tokenization algorithms such as Byte-Pair Encoding (BPE) and SentencePiece are trained predominantly on modern language corpora and produce inefficient segmentations when applied to classical Indian languages.
By Poornima Kumaresan, Pavithra Muruganantham, Lakshmi Rajendran, Santhosh Sivasubramani
arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.
By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv:2608.23120v1 Announce Type: cross
Abstract: Pnar, an Austroasiatic language spoken by approximately 0.4 million people in the Jaintia Hills of Meghalaya, lacks the digital corpora and natural l...
By Edawanbiang Dhar Surmila Thokchom, Thoudam Doren Singh
arXiv:2609.17539v1 Announce Type: new
Abstract: We present MudawanSn, a gold-standard resource of 1,271 sentence-aligned pairs manually translated from Wolof into Modern Standard Arabic (MSA). The so...
By Mouhamed Mbaye, Thierno Diop
arXiv:2606.13218v2 Announce Type: replace
Abstract: Arabic and Hebrew, as closely related Semitic languages, share many words with similar surface forms, including true cognates, false friends, and m...
By Junhong Liang, Noor Abo Mokh, Bashar Alhafni
arXiv:2011.03783v3 Announce Type: replace-cross
Abstract: In this work, we introduce the construction of a machine translation (MT) assisted and human-in-the-loop multilingual parallel corpus with an...
By Lifeng Han, Najet Hadj Mohamed, Malak Rassem, Gareth Jones, Alan Smeaton, Goran Nenadic
arXiv:2608.30092v1 Announce Type: cross
Abstract: We present Arkios, a 1.04B-parameter dense transformer pretrained from scratch on 150B tokens of bilingual English-Nepali text, using a custom single...
By Sajal Regmi, Siddhartha Pudasaini, Chetan Phakami Pun
The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.
By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
arXiv:2608. 05850v1 Announce Type: cross Abstract: We present MameLoshnLM, the first open-source 8B-parameter language model built specifically for Yiddish.
By Uri Katz, Omer Goldman, Tomasz Limisiewicz, Reut Tsarfaty, Noah A. Smith
The paper presents a corpus of 1,262 Classical Tamil verse‑commentary pairs and evaluates several neural representation learning models—including recurrent, Transformer, Siamese, mBART‑style encoder‑decoder, and decoder‑only language models—against a TF‑IDF baseline. Experiments reveal limited gains: token‑F1 scores range from 0.02 to 0.20, the encoder‑decoder continues to lower training loss even after validation loss rises, and the decoder‑only model only reproduces authentic word order in 95.5% of minimal‑pair tests but fails to generate held‑out commentary content. The authors release the extraction and evaluation protocol while noting that redistribution of the source commentaries requires permission.
By Amrit Gopinath, Sangeetha Sivanesan