arXiv Machine Learning
Aug 31

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

The paper presents an automated pipeline that generates high‑quality Quranic datasets, including 848 hours of audio and 286,000 annotated utterances, by collecting recitations, segmenting at pause points with a fine‑tuned wav2vec2‑BERT model, transcribing segments, and verifying transcripts using a novel Tasmeea algorithm. It introduces qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules, and a custom Quran Phonetic Script (QPS) for encoding Tajweed. A multi‑level CTC model trained on this data achieves a 0.21% phoneme error rate on the test set and 1.94% on qdat_bench, with a 75.8% Tajweed F1 score.

By Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas
arXiv Computation and Language
Sep 21

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

arXiv:2609.22038v1 Announce Type: new Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Qura...

By Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
arXiv Computation and Language
Sep 14

What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts

The paper presents a human‑annotated dataset of 100 Quran recitation recordings, identifying 348 scored units and 162 localized events across ten combined labels. An evaluator that scores both labels and word positions achieves a label‑aware F1 of 0.525 and a localization F1 of 0.826, with adapted production cleaner/alignment components yielding similar scores. A pilot experiment with eight 20‑minute runs across three coders and eight models shows wide variance in F1 (0.143–0.892) and highlights that most gold events are detected, leaving only span extent and label conventions as the remaining challenges.

By Mohamad Al Mdfaa, Nursultan Askarbekuly, Ahmed Helaly, Ubai Sandouk, Manuel Mazzara
arXiv AI
Jun 19

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.

By Nabil Mosharraf Hossain (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, United Kingdom), Unaizah Obaidellah (University of Malaya, Malaysia)
arXiv Computation and Language
Aug 28

AraMS-28k: The Largest Publicly Released Line-Level Dataset of Historical Arabic Manuscripts with Margin and Insertion-Anchor Annotations

AraMS-28k is the largest publicly released line‑level dataset of genuine historical Arabic manuscripts, containing 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main‑text and 629 margin). The dataset spans three script traditions—Naskh, Ruq'ah, and Maghrebi—and includes a lithographed printed edition for format diversity. Each line is labeled as main‑text or margin, with margin lines that have a clear attachment point annotated with an insertion anchor to recover the manuscript’s true non‑linear reading order; both fully vocalized and diacritic‑normalized transcriptions are provided, and the data was produced via the RefLAM pipeline combining OCR, clean transcriptions, and human review. "whyItMatters":"The dataset’s comprehensive line‑level annotations, including reading‑order anchors and dual transcription formats, enable reproducible research on Arabic manuscript recognition, layout analysis, and reading‑order recovery under a CC BY‑NC‑SA 4.0 license."

By Mohamed Guechaoui, Mohamed Diaa Zellagui, Souleyman Chaib, Sahraoui Dhelim