arXiv Computation and Language By Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

Read the original on arXiv Computation and Language →

The Flow has not summarised this story yet — read it at arXiv Computation and Language.

arXiv AI
Aug 24

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.

By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
Hugging Face Trending Papers
Jul 22

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.

arXiv Computation and Language
Aug 28

Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap

The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.

By Salima Lamsiyah, Ruslan Mitkov
arXiv Computation and Language
3d ago

Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

HalluScoring 2026 is a shared task that evaluates hallucination detection and factual verification in Arabic question answering, focusing on generalization to unseen questions and LLMs. It comprises two main tasks with four subtasks: binary hallucination detection (Subtasks 1.1 and 1.2) and answer verification against six candidates in Islamic and general knowledge domains (Subtasks 2.1 and 2.2). Thirteen teams participated, with the top system achieving AUC‑ROC scores of 0.772 and 0.767 for detection, and 0.882 and 0.857 for verification.

By Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman
arXiv Machine Learning
Aug 31

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

The paper presents an automated pipeline that generates high‑quality Quranic datasets, including 848 hours of audio and 286,000 annotated utterances, by collecting recitations, segmenting at pause points with a fine‑tuned wav2vec2‑BERT model, transcribing segments, and verifying transcripts using a novel Tasmeea algorithm. It introduces qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules, and a custom Quran Phonetic Script (QPS) for encoding Tajweed. A multi‑level CTC model trained on this data achieves a 0.21% phoneme error rate on the test set and 1.94% on qdat_bench, with a 75.8% Tajweed F1 score.

By Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas