Hugging Face Trending Papers

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.

arXiv Computation and Language
3d ago

Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

HalluScoring 2026 is a shared task that evaluates hallucination detection and factual verification in Arabic question answering, focusing on generalization to unseen questions and LLMs. It comprises two main tasks with four subtasks: binary hallucination detection (Subtasks 1.1 and 1.2) and answer verification against six candidates in Islamic and general knowledge domains (Subtasks 2.1 and 2.2). Thirteen teams participated, with the top system achieving AUC‑ROC scores of 0.772 and 0.767 for detection, and 0.882 and 0.857 for verification.

By Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman
arXiv AI
Sep 1

Beyond Fluency: A Rubric-Based Benchmark for Evaluating Saudi Dialect and Cultural Competence in Large Language Models

The paper introduces a rubric-based benchmark to evaluate Saudi Arabic dialect and cultural competence in large language models. It comprises 31 expert-authored prompts covering idiomatic, pragmatic, lexical, and culturally embedded aspects, each paired with an expert-established ground truth. Four state-of-the-art models were scored, revealing that none exceeded 55% accuracy and that ambiguous framing was the most common error type.

By Ghassan Al-Sumaidaee, Sajjad Abdoli, Ahmed Rashad, Maxim Legg
arXiv Computation and Language
Sep 23

ARAFA: An LLM-Generated Arabic Fact-Checking Dataset

A new large-scale Arabic fact‑checking dataset called Arafa has been created using an automated pipeline that generates claims from Arabic Wikipedia, mutates them into counterfactuals, and validates them against supporting or refuting evidence. The dataset contains 181,976 claim‑evidence pairs labeled as supported, refuted, or not enough information, and human evaluation shows high inter‑annotator agreement and strong validation accuracy. Fine‑tuned transformer models on Arafa achieve a Macro F1‑score of 77%, demonstrating its usefulness for Arabic fact‑checking tasks.

By Christophe Khalil, Shady Elbassuoni, Rida Assaf
arXiv Computation and Language
Sep 21

QuranicMMLU: A Cognitively-Aware Benchmark for Evaluating Generative AI Solutions on Quranic Linguistic Knowledge

arXiv:2609.22038v1 Announce Type: new Abstract: We introduce QuranicMMLU, a benchmark for evaluating generative AI on Quranic Arabic across multiple dimensions of linguistic complexity. Existing Qura...

By Rawan El Ghali, Umm Kulsoom, Anas Madkoor, Dima Faris Alsaudi, Roaa Abdelmagid, Roaa Ibrahim, Raghad Mousa, Hamza Aljaji, Abdullah Khanafer, Abdallah Alkanani, Salah Feras Alali, Rawan Khaled Mohamed, Ehsaneddin Asgari
arXiv AI
Sep 25

PROOF: Profiling Reliability of Object-Level Facts in Large Language Models

PROOF is a benchmark that profiles the reliability of object-level facts in instruction-tuned language models by converting a frozen Wikidata snapshot into 18,486 English multiple-choice questions grounded in 11,779 semantic facts across 101 classes, 392 properties, and 14 domains. Each question includes an explicit "I don't know" option, a "No correct option" control, and nine controlled formulations, with 1,849 questions designed as no-correct-option traps. The study evaluates 18 open-weight model deployments on 166,374 prompts, revealing wide variability in factual accuracy, sensitivity to wording changes, and the impact of decoder perturbations.

By Andrei Chetvergov, Mikhail Solovev, Timofei Sivoraksha, Stepan Ukolov, Valeriia Kuschenko, Alexander Evseev, Sergey Bolovtsov
arXiv AI
Sep 2

EDRAC: Benchmarking Arabic Dialect Reading Comprehension

EDRAC is the first large‑scale benchmark for dialectal Arabic machine reading comprehension and generative question answering, covering five major dialects—Egyptian, Moroccan, Emirati, Syrian, and Saudi. It contains 499 passages from naturally spoken interactions and 4,977 QA pairs produced via a human–LLM collaborative pipeline. The benchmark evaluates Arabic‑centric and multilingual large language models, revealing gaps between semantic answer quality and dialectal fidelity and underscoring limitations of current evaluation metrics for dialectal Arabic generation.

By Noor Abo Mokh, Kirill Chirkunov, Teresa Lynn, Nizar Habash, Reham Marzouk, Malik H. Altakrori, Younes Samih, Muhammed Abu Odeh, Nour Rabih, Rahaf Alshahrani, Hamad Alshehhi, Hamdan Al-Ali, Muhra Almahri, Besher Hassan, Mohamed Anwar, Abed Alhakim Freihat, Preslav Nakov, Alham Fikri Aji
arXiv AI
Aug 24

Ansari: A Retrieval-Grounded Islamic AI Assistant -- Architecture, Deployment, and Lessons from 140,000 Conversations

Ansari is a retrieval‑grounded Islamic AI assistant that has handled over 140,000 conversations in more than 25 languages since June 2023. It uses an agentic retrieval loop where a language model searches authenticated Islamic corpora—including the Qur’an, hadith collections, fiqh encyclopedias, and tafsir sources—and answers only based on retrieved content, providing citations for verification. The paper details Ansari’s architecture, multi‑platform deployment, evaluation results (including top performance on the IslamicMMLU leaderboard and strong resistance to false premises), and lessons for faith‑sensitive LLM deployments.

By M Waleed Kadous, Amr Elsayed, Abdullah Al Nahas, Ashraf Haress
arXiv Computation and Language
Sep 7

ConfRAG: Confidence-Guided Retrieval-Augmenting Generation

ConfRAG introduces a confidence-guided approach to reduce hallucinations in large language models and selectively trigger Retrieval-Augmented Generation (RAG) only when the model is uncertain. The ConfQA fine‑tuning strategy trains the model to answer correctly or respond with "I am unsure," achieving a drop in hallucination rates from 20‑40% to below 5% across factuality benchmarks. Building on ConfQA, ConfRAG limits external retrievals by more than 30% while maintaining over 95% accuracy in ideal scenarios.

By Yin Huang, Yifan Ethan Xu, Kai Sun, Vera Yan, Alicia Sun, Haidar Khan, Jimmy Nguyen, Jingxiang Chen, Mohammad Kachuee, Zhaojiang Lin, Yue Liu, Aaron Colak, Anuj Kumar, Wen-tau Yih, Xin Luna Dong