Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.
EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.
By Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
The paper introduces the Middle East Cultural Sensitivity Score (MECSS) to quantify Orientalist bias in large language models, converting Said’s seven Orientalist operations into measurable dimensions. Using 280 conversations, it finds that GPT‑4 and Falcon3‑7B‑Instruct systematically reproduce Orientalist patterns, with Falcon scoring higher despite being regionally built. The study highlights that geographic origin alone does not mitigate bias and identifies a new failure mode, "Said‑washing," present in 87.9% of GPT‑4 interactions.
By Maha Shahid
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
arXiv:2609.11334v1 Announce Type: cross
Abstract: Natural Language Inference processes pairs of sentences to extract their semantic relations. NLI has been a hot research topic, integrated as a main...
By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
arXiv:2606. 20255v1 Announce Type: cross Abstract: We introduce the Meaning Intelligence Framework (MIF), a nine-dimension annotation and evaluation schema for Nigerian public discourse that separates surface sentiment from true communicative intent.
By Celestine Achi
arXiv:2603. 13891v2 Announce Type: replace-cross Abstract: Large language models (LLMs) are increasingly used for automated text annotation in tasks ranging from academic research to content moderation and hiring.
By Petter T\"ornberg
arXiv:2608.01291v2 Announce Type: replace
Abstract: We present ArabicDialectSafety, a human-curated Arabic safety dataset of 25,071 prompts covering six Arabic varieties: Modern Standard Arabic, Syri...
By Wajdi Zaghouani, Md. Rafiul Biswas, Kholoud Khalil Aldous, Mabrouka Bessghaier
arXiv:2605. 25420v2 Announce Type: replace-cross Abstract: Large language model safety evaluation remains heavily English-centered, leaving low-resource languages under-measured even when models are deployed globally.
By Khalid Yusuf Dahir
arXiv:2606. 01322v1 Announce Type: cross Abstract: Safety evaluation of Large Language Models (LLMs) remains heavily English-centric, leaving Low-Resource Languages (LRLs), particularly African ones, critically underexplored.
By Victor Akinode, Senyu Li, Wassim Hamidouche, Waqas Zamir, Inbal Becker-Reshef, David Ifeoluwa Adelani
arXiv:2607. 20056v1 Announce Type: cross Abstract: Aspect-based sentiment analysis (ABSA) in Arabic must recover both explicitly stated aspects and implicit aspects that are never named in the text.
By Lujain A. Alawwad
arXiv:2608.21985v1 Announce Type: new
Abstract: As the adoption of large language models (LLMs) grows in Arabic-speaking regions, ensuring their safety and cultural alignment is increasingly critical...
By Fidaa Abed, Haidar Khan, M Saiful Bari, Babar Khan, Abdalghani Abujabal