arXiv:2405.06818v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) for Ghana's 73 living indigenous languages remains deeply fragmented, under-resourced, and heavily skewed toward...
By Sheriff Issaka, Erick Rosas Gonzalez, Colene Agbo, Evans Kofi Agyei, Shruti Tyagi, John Emeka Eze, Enock Appiah Tieku, Junlin Fang, Thanh Do Nguyen, Juliet Arthur, Zhaoyi Zhang, Mihir Heda, Keyi Wang, Yinka Ajibola, Rebecca Akpanglo-Nartey, Frank Lawrence Nii Adoquaye Acquaye, Dennis Owusu, Jerry John Kponyo, Stephen Moore, Isaac Wiafe, Sean Du
arXiv:2601.09716v2 Announce Type: replace
Abstract: Natural Language Processing (NLP) is rapidly transforming research methodologies across disciplines, yet African languages remain largely underrepr...
By Derguene Mbaye, Tatiana D. P. Mbengue, Madoune R. Seye, Moussa Diallo, Mamadou L. Ndiaye, Dimitri S. Adjanohoun, Cheikh S. Wade, Djiby Sow, Jean-Claude B. Munyaka, Jerome Chenal
arXiv:2606. 12708v1 Announce Type: cross Abstract: Despite their linguistic diversity and global significance, African languages remain underrepresented in research and resources to support NLP.
By Happy Buzaaba, Cheikh Mouhamadou Bamba Dione, David Ifeoluwa Adelani, Sylvain Kahane, Kim Gerdes, Bruno Guillaume, Kevin Guan, Aremu Anuoluwapo, Naome A. Etori, Shamsuddeen Hassan Muhammad, Utitofon Inyang, Peter Nabende, David Sabiiti Bamutura, Andiswa Bukula, Chinedu Uchechukwu, Rooweither Mabuya, Idris Akinade, Christiane Fellbaum
arXiv:2606. 20255v1 Announce Type: cross Abstract: We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent.
By Celestine Achi
arXiv:2608. 07629v1 Announce Type: cross Abstract: Multilingual neural machine translation models such as NLLB-200 cover 200 languages but leave thousands unsupported, including most Grassfields Bantu languages of Cameroon.
By Samiratu Ntohsi, Neza David Tuyishimire, Anesu Kafesu, Marvin Ogore, Samuel Oluwajunwonlo Babalola, Oche Ankeli
arXiv:2605. 02608v2 Announce Type: replace-cross Abstract: Transformer-based models achieve state-of-the-art dependency parsing for high-resource languages, yet their advantage over simpler architectures in low-resource settings remains poorly understood.
By Kevin Guan, Happy Buzaaba, Christiane Fellbaum
arXiv:2608.21369v1 Announce Type: cross
Abstract: Nigerian Pidgin is one of Africa's most widely spoken languages, yet remains severely underrepresented in language model evaluation. Existing benchma...
By Stephanie Okoye
The paper compares generative and encoder-based neural models for multilingual Named Entity Recognition (NER) across the eleven languages of the Naamapadam benchmark. Five classic model families, four decoder-only large language models fine‑tuned with LoRA and 4‑bit NF4 quantisation, and nine generative models in zero‑to‑5‑shot inference were evaluated under strict CoNLL span‑level metrics. Encoder-based models (mBERT and XLM‑R) achieved substantially higher F1 scores—up to 0.675 on Hindi—than any generative architecture, with gaps of 7.5–40 percentage points; the best few‑shot result reached only 28% of the encoder baseline. The study identifies three language clusters (encoder‑dominant, partial‑coverage, and failure‑zone) and offers deployment guidelines based on transfer learning and low‑resource NLP principles.
By Jakkala Mahesh, Jatavath Shravan Kumar, Komalla Shivani, Sujoy Sarkar
5-Dialects-BN is a new Bangla dialect benchmark that aligns Romanized transliteration with dialectal text, Standard Bangla, English, and subjectivity labels across five regional varieties. The dataset contains 6,000 manually annotated entries from Chittagong, Barisal, Noakhali, Sylhet, and Rangpur, each enriched with five aligned annotations produced and cross‑validated by native speakers and linguistics students. It supports tasks such as dialect identification, normalization, translation, subjectivity classification, and efficient fine‑tuning of multilingual LLMs.
By Md Mahir Jawad, Galib Mahmud Jim, Rafid Ahmed, Mir Sazzat Hossain, Md Fahim, Md Farhad Alam Bhuiyan
arXiv:2609.22494v1 Announce Type: new
Abstract: In recent years, there has been a surge of interest in Cultural NLP, with substantial efforts to create globally inclusive NLP systems. The rapid growt...
By Tania Chakraborty, Eylon Caplan, Zhaoqing Wu, Kevin Cushing, Han Qin, Shreya Havaldar, Dan Goldwasser
The paper surveys language models created for Portuguese, noting that while rapid progress has been made in NLP, development has been uneven across languages. It systematically maps 46 Portuguese models, detailing aspects such as base model, architecture, resources, datasets, licensing, code, data, and weights. The study also traces model evolution phylogenetically, highlights research gaps, and outlines future directions for Portuguese language modeling.
By Jhessica Silva, Carlos Caetano, Helena Maia, Breno Bernard Nicolau de Fran\c{c}a, Sandra Avila, Helio Pedrini
arXiv:2608.04186v3 Announce Type: replace
Abstract: This paper presents a conceptual framework for developing an electronic explanatory dictionary of the Tajik language using large language models (L...
By Mullosharaf K. Arabov, Saidali M. Pirzoda, Behruz A. Sultonov