The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
The paper investigates how transformer-based models and traditional feature-based models capture readability signals across five languages using the ReadMe++ dataset. By applying SHAP to identify key features for traditional classifiers and then probing XLM‑R and language‑specific encoders with TCAV, the authors find that transformers recover surface‑length, syntactic, and lexical‑diversity cues and reflect the CEFR ordinal structure, though alignment varies by model family, language, and layer. The study highlights that high linear separability does not guarantee directional influence, cautioning against overreliance on linear probing for readability features.
By Joshua Wong, Chris Tanner
arXiv:2606. 07706v1 Announce Type: cross Abstract: Vision-Language Models (VLMs) have demonstrated strong performance across multimodal tasks, yet their safety robustness remains an open challenge.
By Rishabh Makwana, Mamta, Deeksha Varshney, Oana Cocarascu
arXiv:2607. 20385v1 Announce Type: cross Abstract: Optical Character Recognition (OCR) for Persian remains substantially less mature than for Latin-script languages despite Persian being spoken by more than 110 million people across multiple countries.
By Pouria Mahdi, Haq Nawaz Malik
arXiv:2606.27314v2 Announce Type: replace
Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...
By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.