AraGenre 2026 is a shared task that challenges participants to perform hierarchical, definition-guided Arabic genre classification. Systems must assign each Arabic text segment both a broad communicative genre and a fine-grained specific genre, using natural-language definitions for 74 unseen specific genres in a zero-shot setting. The released datasets include limited synthetic examples for training and development, while the hidden final benchmark contains noisier, naturally occurring text across Modern Standard Arabic, Classical Arabic, and multiple dialects.
By Mo El-Haj, Saad Ezzini, Shadi Abudalfa, Mustafa Jarrar, Nguyen Minh Chi, Nguyen Minh Quan
NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification shared task series and the second focused on multidialectal Arabic speech processing. It includes five main tasks—Automatic Speech Recognition, Spoken Dialect Identification, Text-to-Speech, Spoken Language Translation, and Spoken Language Understanding—along with eight subtasks that test realistic scenarios such as low‑bandwidth, mixed dialects, code‑switching, out‑of‑domain, and zero‑shot settings. The event attracted 21 teams from at least 13 countries, with 48 test‑phase submissions and 14 system‑description papers, and the results highlight out‑of‑domain generalization as a major bottleneck while showcasing the strengths of Arabic‑specialized speech models, multimodal dialect identification, and ensemble methods.
By Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Est\`eve, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed
arXiv:2608.01291v2 Announce Type: replace
Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syri...
By Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous, Mabrouka Bessghaier
arXiv:2610.09756v1 Announce Type: new
Abstract: Classical Arabic poetry generation requires simultaneously satisfying semantic, linguistic, and fine-grained prosodic constraints. Existing systems typ...
By Ahmad Abbas, Tamara Fakih, Nour Fakih, Ammar Mohanna
arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.
By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov
Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.
By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
arXiv:2601. 12494v3 Announce Type: replace-cross Abstract: Audio large language models (LLMs) enable unified speech understanding and generation, but adapting them to linguistically complex and dialect-rich settings such as Arabic-English remains challenging.
By Hunzalah Hassan Bhatti, Firoj Alam, Shammur Absar Chowdhury
arXiv:2506. 10292v2 Announce Type: replace-cross Abstract: Training deep learning networks with minimal supervision has gained significant research attention due to its potential to reduce reliance on extensive labelled data.
By Ali Almutairi, Abdullah Alsuhaibani, Shoaib Jameel, Aditya Joshi, Gelareh Mohammadi, Imran Razzak
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
arXiv:2609.22796v1 Announce Type: new
Abstract: Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective t...
By Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy, Saad Ezzini, Mo El-Haj, Mustafa Jarrar, Zaid Alyafeai, Bernard Ghanem, Muhammad Abdul-Mageed
A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.
By Christophe Khalil, Shady Elbassuoni, Rida Assaf
The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.
By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg