arXiv AI

WASIL: In-the-Wild Arabic Spoken Interactions with LLMs

arXiv Computation and Language
Sep 11

Nuha-Speech: Building General-Purpose Arabic Speech-LLMs

Nuha‑Speech is a new initiative aimed at creating general‑purpose Arabic speech‑large language models (speech‑LLMs). It includes the construction of a large Arabic Speech Question‑Answering corpus with over 1.5 million samples for instruction tuning, supervised fine‑tuning of Qwen‑Omni model variants at various scales, and a systematic evaluation framework with diverse tasks and tailored metrics. The project seeks to establish foundational infrastructure for Arabic speech‑LLMs amid limited Arabic speech resources.

By Yingzhi Wang, Reem Alhazzani, Muhammad Alqurishi
arXiv AI
Sep 2

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.

By Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
arXiv AI
Sep 21

MENASpeechBank: A Reference Voice Bank with Persona-Conditioned Multi-Turn Conversations for AudioLLMs

MENASpeechBank is a new reference voice bank that provides about 18,000 high‑quality utterances from 124 speakers across multiple MENA countries, covering English, Modern Standard Arabic, and regional Arabic varieties. The dataset is built through a controllable pipeline that creates persona profiles inspired by the World Values Survey, defines a taxonomy of roughly 5,000 conversational scenarios, matches personas to scenarios via semantic similarity, and generates around 417,000 role‑play conversations using an LLM. Synthetic speaker‑conditioned user‑turn audio is produced from reference recordings to maintain speaker diversity, and both synthetic and human‑recorded conversations are evaluated and analyzed for quality.

By Zien Sheikh Ali, Hunzalah Hassan Bhatti, Rabindra Nath Nandi, Shammur Absar Chowdhury, Firoj Alam
arXiv AI
Aug 25

Bulbul: A Dataset for Dialectal Arabic Speech Recognition

arXiv:2608.21950v1 Announce Type: cross Abstract: Arabic automatic speech recognition (ASR) faces unique challenges due to diglossia, extensive regional dialect variation, and limited speech resource...

By Ahmed Ashraf, Aisha Alansari, Fadel Al Abbas, Nada Almarwani, Samah Aloufi, Saad Ezzini, Maged S. Al-Shaibani, Doaa Dalaq, AbdelRahim A. Elmadany, Muhammad Abdul-Mageed, Mohamed Mehdi Trigui, Dania Refai, Layan Refai, Mohamed Akrout, Mustafa Jarrar, Wasfi G. Al-Khatib, Alaa Dalaq, Darin El-Nakla, Samir Abdaljalil, Abdulrahman Al-Fakih, Nour El Imane Zeghib, Moussa Redah, Salmane Chafik, Mohamed El-Attar, Rima Grati, Sarah Kohail, Malak Alkhorasani, Khadijah Al Safwan, Ismail M. Mudhaffar, Ali Altam, Ahmed Al-Shaikh, Adnan Saeed, Hamzah Luqman
arXiv Computation and Language
Sep 24

NADI 2026: The Second Multidialectal Arabic Speech Processing Shared Task

NADI 2026 is the seventh edition of the Nuanced Arabic Dialect Identification shared task series and the second focused on multidialectal Arabic speech processing. It includes five main tasks—Automatic Speech Recognition, Spoken Dialect Identification, Text-to-Speech, Spoken Language Translation, and Spoken Language Understanding—along with eight subtasks that test realistic scenarios such as low‑bandwidth, mixed dialects, code‑switching, out‑of‑domain, and zero‑shot settings. The event attracted 21 teams from at least 13 countries, with 48 test‑phase submissions and 14 system‑description papers, and the results highlight out‑of‑domain generalization as a major bottleneck while showcasing the strengths of Arabic‑specialized speech models, multimodal dialect identification, and ensemble methods.

By Peter Sullivan, Bashar Talafha, Ahmed Ashraf, Fethi Bougares, Haroun Elleuch, Chiyu Zhang, AbdelRahim Elmadany, Youssef Mohamed, Salima Mdhaffar, Yannick Est\`eve, Mohamed Elhoseiny, Hamzah Luqman, Nizar Habash, Muhammad Abdul-Mageed
arXiv Computation and Language
Aug 28

Scaling phoneme-based TTS augmentation for ASR: A unified pipeline and controlled study

The paper introduces a unified phoneme‑based TTS‑to‑ASR augmentation pipeline that uses a multilingual TTS model with language‑ID conditioning and incorporates grapheme‑to‑phoneme conversion, reference‑speech filtering, and candidate‑text selection. It proposes phoneme‑frequency‑guided selection (PFGS) to rank sentences based on phoneme frequencies from real ASR labels, and demonstrates that random augmentation and PFGS both improve ASR performance across Arabic, French, Italian, and Portuguese test sets, with PFGS yielding up to a 19.3% relative WER reduction. The study also shows that filtering reference speech can further lower WER by up to 0.59 points on certain datasets.

By Zhen Wang, TianRui Wu, RongQi Han, Hao Wu, Wei Liang
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv Computation and Language
Sep 22

AlexandriaX 2026: The First Shared Task on Dialectal Arabic Machine Translation

arXiv:2609.22796v1 Announce Type: new Abstract: Dialectal Arabic machine translation (MT) remains challenging despite recent progress in Arabic language technologies, particularly because effective t...

By Abdellah El Mekki, AbdelRahim A. Elmadany, Samar M. Magdy, Saad Ezzini, Mo El-Haj, Mustafa Jarrar, Zaid Alyafeai, Bernard Ghanem, Muhammad Abdul-Mageed