arXiv AI

When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages

The paper "When Explanations Cannot Be Read: Measuring and Correcting SHAP and LIME Rendering for Right-to-Left Languages" identifies that standard SHAP and LIME visualizations, designed for left‑to‑right scripts, fail to display attribution values correctly for right‑to‑left languages such as Urdu, Arabic, Persian, and Hebrew. It introduces SHAP‑RTL, a rendering layer that preserves the original attribution values while correcting reading direction, script shaping, and font selection for each language. The authors evaluate SHAP‑RTL on hate‑and‑offensive‑language datasets using TF‑IDF and logistic regression, showing that default rendering yields high character error rates, while SHAP‑RTL maintains correct visualizations across all tested languages.

arXiv Computation and Language
Aug 28

Why Current XAI Is Not Enough for Arabic NLP: A Critical Survey of the Explainability Gap

The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.

By Salima Lamsiyah, Ruslan Mitkov
arXiv Computation and Language
Sep 11

Analyzing Traditional and Neural Approaches to Multilingual Readability Assessment

The paper investigates how transformer-based models and traditional feature-based models capture readability signals across five languages using the ReadMe++ dataset. By applying SHAP to identify key features for traditional classifiers and then probing XLM‑R and language‑specific encoders with TCAV, the authors find that transformers recover surface‑length, syntactic, and lexical‑diversity cues and reflect the CEFR ordinal structure, though alignment varies by model family, language, and layer. The study highlights that high linear separability does not guarantee directional influence, cautioning against overreliance on linear probing for readability features.

By Joshua Wong, Chris Tanner
arXiv Computation and Language
Sep 15

Beyond Surface Forms: A Comprehensive, Mechanism-Oriented Taxonomy of Indirect Linguistic Encoding for LLM-Based Coded Language Detection

arXiv:2606.27314v2 Announce Type: replace Abstract: To avoid moderation and surveillance on social media, some users routinely invent indirect linguistic expressions (ILE) that camouflage sensitive m...

By Hamid Reza Firoozfar, Mohammadsadegh Abolhasani, Reza Mousavi, Paul Jen-Hwa Hu
Hugging Face Trending Papers
Jul 22

HalluTruthQA: A Fine-Grained Benchmark for Hallucination Detection, Localization, and Explanation in Arabic Question Answering

Large language models (LLMs) can generate fluent Arabic answers, yet factual errors remain difficult to detect, localize, explain, and verify. Existing hallucination benchmarks often provide response-level labels, with limited support for identifying the exact erroneous content, explaining why it is incorrect, or selecting the correct factual answer.

arXiv Machine Learning
Sep 24

ASCIIBench: Evaluating Language-Model-Based Understanding of Visually-Oriented Text

ASCIIBench is a new benchmark that evaluates large language models on generating and classifying ASCII-text images, using a dataset of 5,315 labeled ASCII images. The authors also release a fine‑tuned CLIP model adapted to capture ASCII structure for evaluation. Their analysis shows that cosine similarity on CLIP embeddings fails to separate most categories, indicating a representation bottleneck rather than generational variance.

By Kerry Luo, Michael Fu, Joshua Peguero, Husnain Malik, Anvay Patil, Joyce Lin, Megan Van Overborg, Ryan Sarmiento, Kevin Zhu
arXiv AI
Sep 4

IndicSafeEval: Safety Robustness of Large Language Models under Multilingual Persuasive Jailbreak Attacks

IndicSafeEval is a new evaluation framework that tests the safety robustness of large language models against persuasion-based jailbreak attacks in Indian languages. The benchmark includes 7,200 adversarial prompts covering ten safety-critical content categories, six persuasive strategies, and four languages (Hindi, Bengali, Marathi, Punjabi). Experiments show that model safety varies significantly across languages, prompt styles, and risk categories, revealing that current English-centric safety evaluations miss important multilingual vulnerabilities.

By Saikat Mondal, Mamta, Deeksha Varshney, Oana Cocarascu, Asif Ekbal
arXiv Computation and Language
Sep 16

AraMIP: Extending MIPVU Towards Metaphor Identification in Arabic

AraMIP introduces a new guideline for annotating metaphors in Arabic, building upon the established MIPVU framework and tailoring it to Arabic’s linguistic features. The authors distinguish three figurative types—Isti'ara (metaphor), kinaya (metonymy/indirect expression), and tashbih (simile)—and apply the procedure to a pilot dataset of 300 sentences (5,277 words). Their analysis highlights Arabic‑specific challenges such as morphological complexity, inconsistent dictionary sense ordering, and a lack of standardized contextual materials for annotators.

By Mandar Marathe, Manar Ali, Sara Nabhani, Raia Abu Ahmad, Ibrahim Baroud, Omar Momen
arXiv Computer Vision
Sep 14

ChitraMiti: Benchmarking Visual Grounding and Modality Reliance in Bengali Geometric Reasoning

ChitraMiti introduces a synthetic benchmark of 12,874 Bengali planar geometry problems with structured 15‑attribute descriptions, alongside a complementary set of 500 textbook diagrams. Using a three‑phase protocol that tests diagram‑only, diagram‑plus‑description, and description‑only inputs, the study finds that description‑only performance matches diagram‑plus‑description performance across several VLMs, yet models still struggle with cross‑modal verification. Fine‑tuning on ChitraMiti improves results on both datasets, though a gap remains compared to the best zero‑shot model.

By Khan Raiyan Ibne Reza, Sanjana Aktar Maria, Sumaiya Tabassum Nimi, Md Adnan Arefeen
arXiv AI
Aug 17

Jais 2: A Family of Arabic-Centric Open Large Language Models

arXiv:2608. 13580v1 Announce Type: cross Abstract: Jais 2 is a family of Arabic-centric large language models developed jointly by MBZUAI, Cerebras, and Inception, designed to advance Arabic-centric language modeling, with strong performance across the Arabic and culturally grounded benchmarks evaluated in this report.

By Mohamed Anwar, Abed Alhakim Freihat, George Ibrahim, Mostafa Awad, Abdelrahman Sadallah, Gurpreet Gosal, Gokulakrishnan Ramakrishnan, Sarath Chandran, Biswajit Mishra, Rituraj Joshi, Ahmed Frikha, Etienne Goffinet, Abhishek Maiti, Ali El Filali, Sarah AlBarri, Samujjwal Ghosh, Rahul Pal, Parvez Mullah, Awantika Shukla, Sajid siddiki, Samta Kamboj, Onkar Pandit, Sunil Kumar Sahu, AbdelRahman Elbadawy, Amr Mohamed, Ahmad Chamma, Evan Dufraisse, Abdelaziz Bounhar, Dani Bouch, Hadi Abdine, Guokan Shang, Fajri Koto, Yuxia Wang, Zhuohan Xie, Ali Mekky, Rania Elbadry, Sarfraz Ahmad, Momina Ahsan, Omar El Herraoui, Daniil Orel, Hasan Iqbal, Kareem Elzeky, Mervat Abassy, Kareem Elozeiri, Saadeldine Eletter, Farah Atif, Nurdaulet Mukhituly, Haonan Li, Xudong Han, Aaryamonvikram Singh, Zainul Abedien Ahmed Quraishi, Neha Sengupta, Larry Murray, Avraham Sheinin, Joel Hestness, Natalia Vassilieva, Hector Xuguang Ren, Zhengzhong Liu, Michalis Vazirgiannis, Preslav Nakov