HalluScoring 2026 is a shared task that evaluates hallucination detection and factual verification in Arabic question answering, focusing on generalization to unseen questions and LLMs. It comprises two main tasks with four subtasks: binary hallucination detection (Subtasks 1.1 and 1.2) and answer verification against six candidates in Islamic and general knowledge domains (Subtasks 2.1 and 2.2). Thirteen teams participated, with the top system achieving AUC‑ROC scores of 0.772 and 0.767 for detection, and 0.882 and 0.857 for verification.
By Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman
arXiv:2605. 31483v1 Announce Type: cross Abstract: Despite Bengali being the sixth most spoken language in the world, no prior work has systematically evaluated hallucination in large language models (LLMs) for Bengali.
By Shefayat E Shams Adib, Ahmed Alfey Sani, Ekramul Alam Esham, Ajwad Abrar, Ishmam Tashdeed, Md Taukir Azam Chowdhury
The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.
By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.
By Christophe Khalil, Shady Elbassuoni, Rida Assaf
arXiv:2609.22038v1 Announce Type: new
Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Qura...
By Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.
By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov