Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.
By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
arXiv:2609.14967v1 Announce Type: new
Abstract: Quranic text is distributed in two orthographic forms that are byte-level distinct: the Uthmani script used in every printed mushaf, and the Standard (...
By Yahya Mohamed Elnawasany
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
HalluScoring 2026 is a shared task that evaluates hallucination detection and factual verification in Arabic question answering, focusing on generalization to unseen questions and LLMs. It comprises two main tasks with four subtasks: binary hallucination detection (Subtasks 1.1 and 1.2) and answer verification against six candidates in Islamic and general knowledge domains (Subtasks 2.1 and 2.2). Thirteen teams participated, with the top system achieving AUC‑ROC scores of 0.772 and 0.767 for detection, and 0.882 and 0.857 for verification.
By Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman
The paper presents an automated pipeline that generates high‑quality Quranic datasets, including 848 hours of audio and 286,000 annotated utterances, by collecting recitations, segmenting at pause points with a fine‑tuned wav2vec2‑BERT model, transcribing segments, and verifying transcripts using a novel Tasmeea algorithm. It introduces qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules, and a custom Quran Phonetic Script (QPS) for encoding Tajweed. A multi‑level CTC model trained on this data achieves a 0.21% phoneme error rate on the test set and 1.94% on qdat_bench, with a 75.8% Tajweed F1 score.
By Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas
EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.
By Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
arXiv:2609.16006v1 Announce Type: cross
Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test wha...
By Enes Altinisik, Hamdy Mubarak, Masoomali Fatehkia, Husrev_Taha_Sencar Husrev Taha Sencar
The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.
By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
The paper presents a human‑annotated dataset of 100 Quran recitation recordings, identifying 348 scored units and 162 localized events across ten combined labels. An evaluator that scores both labels and word positions achieves a label‑aware F1 of 0.525 and a localization F1 of 0.826, with adapted production cleaner/alignment components yielding similar scores. A pilot experiment with eight 20‑minute runs across three coders and eight models shows wide variance in F1 (0.143–0.892) and highlights that most gold events are detected, leaving only span extent and label conventions as the remaining challenges.
By Mohamad Al Mdfaa, Nursultan Askarbekuly, Ahmed Helaly, Ubai Sandouk, Manuel Mazzara
E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.
By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
arXiv:2609. 28245v1 Announce Type: cross Abstract: Large language models (LLMs) have shown strong performance in creative text generation, yet their ability to produce culturally grounded and stylistically constrained literary forms remains underexplored.
By AbdulRahman A. Morsy (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States), Aya Zirikly (Department of Computer Science, School of Engineering and Applied Sciences, George Washington University, Washington DC, United States, Center for Speech and Language Processing, Whiting School of Engineering, Johns Hopkins University, Baltimore MD, United States)