A Corpus-Aligned Uthmani-to-Standard Quranic Word Mapping and a Deterministic Recitation Validator
Read the original on arXiv Computation and Language →The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The Flow has not summarised this story yet — read it at arXiv Computation and Language.
The paper presents an automated pipeline that generates high‑quality Quranic datasets, including 848 hours of audio and 286,000 annotated utterances, by collecting recitations, segmenting at pause points with a fine‑tuned wav2vec2‑BERT model, transcribing segments, and verifying transcripts using a novel Tasmeea algorithm. It introduces qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules, and a custom Quran Phonetic Script (QPS) for encoding Tajweed. A multi‑level CTC model trained on this data achieves a 0.21% phoneme error rate on the test set and 1.94% on qdat_bench, with a 75.8% Tajweed F1 score.
arXiv:2609.22038v1 Announce Type: new Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Qura...
The paper presents a human‑annotated dataset of 100 Quran recitation recordings, identifying 348 scored units and 162 localized events across ten combined labels. An evaluator that scores both labels and word positions achieves a label‑aware F1 of 0.525 and a localization F1 of 0.826, with adapted production cleaner/alignment components yielding similar scores. A pilot experiment with eight 20‑minute runs across three coders and eight models shows wide variance in F1 (0.143–0.892) and highlights that most gold events are detected, leaving only span extent and label conventions as the remaining challenges.
arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.
AraMS-28k is the largest publicly released line‑level dataset of genuine historical Arabic manuscripts, containing 14 books, 3,043 pages, and 28,600 annotated text lines (27,971 main‑text and 629 margin). The dataset spans three script traditions—Naskh, Ruq'ah, and Maghrebi—and includes a lithographed printed edition for format diversity. Each line is labeled as main‑text or margin, with margin lines that have a clear attachment point annotated with an insertion anchor to recover the manuscript’s true non‑linear reading order; both fully vocalized and diacritic‑normalized transcriptions are provided, and the data was produced via the RefLAM pipeline combining OCR, clean transcriptions, and human review. "whyItMatters":"The dataset’s comprehensive line‑level annotations, including reading‑order anchors and dual transcription formats, enable reproducible research on Arabic manuscript recognition, layout analysis, and reading‑order recovery under a CC BY‑NC‑SA 4.0 license."
arXiv:2609.17539v1 Announce Type: new Abstract: We present MudawanSn, a gold-standard resource of 1,271 sentence-aligned pairs manually translated from Wolof into Modern Standard Arabic (MSA). The so...