arXiv Computation and Language

What Counts as a Mistake? Annotating Recitation Events in Quran Memorization Transcripts

The paper presents a human‑annotated dataset of 100 Quran recitation recordings, identifying 348 scored units and 162 localized events across ten combined labels. An evaluator that scores both labels and word positions achieves a label‑aware F1 of 0.525 and a localization F1 of 0.826, with adapted production cleaner/alignment components yielding similar scores. A pilot experiment with eight 20‑minute runs across three coders and eight models shows wide variance in F1 (0.143–0.892) and highlights that most gold events are detected, leaving only span extent and label conventions as the remaining challenges.

arXiv Machine Learning
Aug 31

Automatic Pronunciation Error Detection and Correction of the Holy Quran's Learners Using Deep Learning

The paper presents an automated pipeline that generates high‑quality Quranic datasets, including 848 hours of audio and 286,000 annotated utterances, by collecting recitations, segmenting at pause points with a fine‑tuned wav2vec2‑BERT model, transcribing segments, and verifying transcripts using a novel Tasmeea algorithm. It introduces qdat_bench, a benchmark covering phonemes, diacritization, and Tajweed rules, and a custom Quran Phonetic Script (QPS) for encoding Tajweed. A multi‑level CTC model trained on this data achieves a 0.21% phoneme error rate on the test set and 1.94% on qdat_bench, with a 75.8% Tajweed F1 score.

By Abdullah Abdelfattah, Mahmoud I. Khalil, Hazem Abbas
arXiv AI
Jun 19

A Comparative Study of Pretrained Transformer Models for Quranic ASR: Speech Representations, Label Formats, and Dataset Composition

arXiv:2606. 19747v1 Announce Type: new Abstract: Quran Automatic Speech Recognition (ASR) aims to convert Quranic recitation into text, enabling applications such as aided memorisation tools and Quranic search engines.

By Nabil Mosharraf Hossain (Greentech Apps Foundation, United Kingdom), Riasat Islam (Greentech Apps Foundation, United Kingdom, Queen Mary University of London, United Kingdom), Unaizah Obaidellah (University of Malaya, Malaysia)
arXiv Computation and Language
Sep 21

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

arXiv:2609.22038v1 Announce Type: new Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Qura...

By Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
Hugging Face Trending Papers
Jul 22

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.

arXiv Computation and Language
Sep 11

Unadapted Multilingual ASR on a Garrusi Kurdish Evaluation Set: A Common-Reference Staged Normalization Analysis

The paper evaluates a multilingual ASR model (MMS‑1B‑all) on a Garrusi Kurdish dataset using a common‑reference staged normalization approach. By normalizing both reference and hypothesis, the authors show that raw Arabic‑script hypotheses yield a 111.70 % WER, which drops to 97.85 % after folding into a reduced orthography, highlighting the impact of orthographic differences on error measurement. A Southern Kurdish fine‑tuned system performs worse, and residual errors are partly due to scoring‑pipeline limitations rather than recognition failures.

By Hiwa Asadpour
arXiv Computation and Language
Sep 15

An Empirical Analysis of Factual Errors in Human-Written Text and Its Application to Factual Error Detection

The paper presents an empirical study of factual errors in human-written text, focusing on corrections in newspaper articles to build a taxonomy of common mistakes such as kanji misconversions and unit errors. It evaluates large language models’ ability to detect these errors, finding that even advanced models like GPT‑5.4 achieve only a 52% word‑level F1 score on synthetic data, underscoring the difficulty of the task. The work highlights the gap in research on factual error detection in human writing compared to LLM hallucinations.

By Kazuma Iwamoto, Kazumasa Omura, Shotaro Ishihara
arXiv Computer Vision
Aug 25

Does a Modern-Handwriting Warm-Up Help Historical Arabic OCR? A Reproducible, Compute-Matched Evaluation on Muharaf and KHATT

arXiv:2608.22316v1 Announce Type: cross Abstract: Whether an intermediate stage of modern Arabic handwriting helps or hurts historical Arabic HTR is usually decided from one implementation and one co...

By Sumaih Almarshad, Maram Alamri, Dona Aloraini, Fares Altuwaim, AlJawharh AlOtaibi, Reem Alyabis, Rayah Aldawsari