arXiv:2605. 02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood.
By Kevin Guan, Happy Buzaaba, Christiane Fellbaum
We present DONDO, a family of open, permissively licensed automatic speech recognition (ASR) base models for African languages, built on the w2v-BERT 2. 0 self-supervised speech encoder.
arXiv:2609.37883v1 Announce Type: new
Abstract: Pre-trained language models (PLMs) with encoder-based architectures have shown impressive capabilities in zero-shot cross-lingual transfer for various...
By Lalita Lowphansirikul, Attapol Rutherford, Jian Gang Ngui, Sarana Nutanong, Peerat Limkonchotiwat
arXiv:2405.06818v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
The paper introduces AraSEG, a new Arabic sentence segmentation corpus covering eight genres and diverse punctuation and document structures. Experiments using AraSEG evaluate large language models, lightweight encoders, and dependency parser-based models, revealing that lightweight encoders and parser-based models outperform LLMs under the most challenging conditions. The study also shows that increasing training data size and genre diversity eventually saturates performance, that cross‑genre generalization remains difficult, and that accurate sentence segmentation significantly improves downstream dependency parsing.
By Mohammed Elkholy, Khalid N. Elmadani, Nizar Habash, Bashar Alhafni
AfriSwitch is a 61.36‑hour, human‑transcribed benchmark of in‑the‑wild code‑switched speech covering 16 African languages and varieties, annotated with switch‑level English span tags, per‑utterance Code‑Mixing Index (CMI), and switch‑point counts. The corpus reveals that code‑switching behaviour varies widely across languages, with no single metric fully capturing how code‑switched a language is. Benchmarking five open and commercial multilingual ASR systems in a zero‑shot setting shows high word error rates, with the best system averaging 35.93% WER and none dropping below 24% on any language, indicating that Africa‑targeted training rather than model scale or nominal language coverage best predicts performance.
By Gabrial Zencha Ashungafac, Busayo Awobade, Tobi Olatunji
ThaiTrees is a 342‑million‑token corpus of Thai text spanning news, Wikipedia, spoken transcripts, and social media, automatically parsed under the Universal Dependencies framework. The authors provide a reproducible pipeline for cleaning, processing, and parsing the data, and release the resulting CoNLL‑U files and a frequency lexicon in machine‑readable formats. This resource enables researchers to search grammatical relations and study syntactic distributions at scale.
By Attapol T. Rutherford, Papatchol Thientong
arXiv:2606. 03219v1 Announce Type: cross Abstract: African languages have very little labelled data, and it is unclear if augmenting the quantity of annotation data reliably enhances downstream performance.
By Anuj Tiwari, Oluwapelumi Ogunremu, Terry Oko-odion, Jesujuwon Egbewale, Hannah Nwokocha
The paper presents a pipeline that leverages large language models to extract grammatical rules, example sentences, and lexicons from descriptive grammar books, producing synthetic parallel corpora for fine‑tuning machine translation models. Evaluated on three low‑resource languages—Kalamang, Tuatschin, and Mandan—the synthetic data improves translation quality over seed‑data baselines in 75% of configurations for Kalamang and 59% for Tuatschin, achieving up to +8.8 ChrF++ gains. A factorial study across 96 configurations identifies which combinations of target part‑of‑speech, retrieval granularity, and sample volume drive performance gains and where they fail, demonstrating that static linguistic documentation can be repurposed for practical translation tools for severely under‑resourced languages.
By Varun Ghat Ravikumar, Sina Ahmadi, Lena J\"ager, Rico Sennrich
The paper surveys NLP research on Nigeria’s three major low‑resource languages—Hausa, Yoruba, and Igbo—covering over 500 languages spoken by 175 million people. It reviews 293 studies, finding that only 27.6% produced new linguistic resources, indicating a heavy reliance on repurposing existing data. The authors highlight under‑explored challenges such as morphological analysis and diacritic representation, and call for collaborative resource enrichment and community support to advance NaijaNLP and low‑resource NLP more broadly.
By Isa Inuwa-Dutse
arXiv:2606. 24825v1 Announce Type: cross Abstract: Part-of-Speech (POS) tagging is a foundational NLP task underpinning machine translation, information extraction, and syntactic parsing.
By Hariom Ingle, Ronit Ghode, Ishwari Gondkar, Jidnyasa Harad, Raviraj Joshi
The paper evaluates large language models (LLMs) on Arabic morphosyntactic tagging and dependency parsing, a challenging task due to rich morphology and orthographic ambiguity. It compares zero‑shot prompting with retrieval‑based in‑context learning across pre‑tokenized, raw‑text, and cascaded settings, finding that relevant demonstrations significantly boost performance. The best LLMs nearly match supervised systems but need extensive annotated data for demonstrations and high computational resources. All code and data are publicly released.
By Mohamed Adel, Bashar Alhafni, Nizar Habash