The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.
By Ali Almutairi, Gelareh Mohammadi, Imran Razzak, Aditya Joshi
arXiv:2608.30425v1 Announce Type: new
Abstract: Cross-lingual aspect-based sentiment analysis (ABSA) transfers knowledge from a source language with annotated data to a target language, enabling fine...
By Jakub \v{S}m\'{i}d, Pavel P\v{r}ib\'{a}\v{n}, Pavel Kr\'{a}l
E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.
By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
Camellia is a new benchmark that tests cultural bias in large language models (LLMs) across nine Asian languages and six Asian cultures. It contains 19,530 manually annotated entities linked to Asian or Western cultures and 2,173 masked social‑media contexts for these entities. Using Camellia, the authors evaluate four multilingual LLMs on cultural context adaptation, sentiment association, and entity extractive QA, finding that models struggle with cultural adaptation, exhibit differing biases across regions and families, and have difficulty understanding context in some Asian languages.
By Tarek Naous, Anagha Savit, Carlos Rafael Catalan, Geyang Guo, Jaehyeok Lee, Kyungdon Lee, Lheane Marie Dizon, Mengyu Ye, Neel Kothari, Sahajpreet Singh, Sarah Masud, Tanish Patwa, Trung Thanh Tran, Zohaib Khan, Alan Ritter, Tanmoy Chakraborty, Yuki Arase, Keisuke Sakaguchi, JinYeong Bak, Wei Xu
The paper introduces STAR‑Ar, a BERT‑BiLSTM‑CRF model designed for the Daleel 2026 Arabic argument mining shared task. It treats argument discourse unit detection and classification as a token‑level sequence labeling problem, achieving an F1‑score of 72.69 on validation and 73.7 on test data. Analysis shows that models trained only on editorial texts perform worse than those trained on debates, mainly due to the smaller editorial dataset.
By Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler
arXiv:2607. 19243v1 Announce Type: cross Abstract: Although Large Language Models (LLMs) demonstrate remarkable multilingual fluency, their internal knowledge representations remain disproportionately biased toward high-resource languages.
By Alexander Manev
arXiv:2608.00207v2 Announce Type: replace
Abstract: Large Language Models (LLMs) perform strongly in English medical tasks but degrade substantially in Arabic, a gap widely attributed to limited trai...
By Chaimae Abouzahir, Musa Khan, Hala Ali-Hassan, Congbo Ma, Khaled Saleh, Yousra Sadqi, Jihad Mallat, Walid Al-Eisawi, Nizar Habash, Farah E. Shamout
arXiv:2609.36194v1 Announce Type: new
Abstract: Extracted sentiment directions can vary across samples even when downstream sentiment classification remains accurate. To evaluate direction reproducib...
By Muhammad Abdullahi Said, Abass Oguntade, Elisha Komolafe, Babangida Sani, Fatima Muhammad Adam, Muhammad Sammani Sani
arXiv:2609.16006v1 Announce Type: cross
Abstract: Large language models (LLMs) increasingly serve users whose expectations are shaped by their cultural context, yet most cultural evaluations test wha...
By Enes Altinisik, Hamdy Mubarak, Masoomali Fatehkia, Husrev_Taha_Sencar Husrev Taha Sencar
The paper surveys the state of Explainable AI (XAI) in Arabic NLP, highlighting three gaps: a method gap where Arabic XAI relies mainly on limited post‑hoc techniques; a task gap with most work focused on classification tasks and little on generation, retrieval, or dialogue; and a linguistic gap where explanations rarely address Arabic‑specific phenomena such as morphology, dialects, and diglossia. It proposes a taxonomy of tasks, methods, linguistic units, and evaluation practices, and outlines a research agenda for linguistically grounded Arabic XAI.
By Salima Lamsiyah, Ruslan Mitkov
The paper compares Knowledge-Graph Based Augmentation (Graph-RAG) with Retrieval-Augmented Generation (RAG) for answering culturally specific questions. Using the LatamQA dataset, Graph-RAG, built automatically from Wikipedia via KGGen, matches RAG performance and reduces the base LLM’s error by 72% with a standard KG and 78% with a benchmark-aware variant. The approach also transfers zero‑shot to Portuguese, showing multilingual applicability.
By Pablo Poulenard, Yannis Karmim, Valentin Barri\`ere
arXiv:2608. 10444v1 Announce Type: cross Abstract: Large language models (LLMs) have made substantial progress on reasoning tasks that require increasingly long and complex inferential chains.
By Si'an Xie (Beijing University of Posts and Telecommunications), Jiaxun Liu (Peking University), Biao Yang (Kuaishou Technology), Wei Yuan (Kuaishou Technology), Fan Yang (Kuaishou Technology), Tingting Gao (Kuaishou Technology), Ming Wu (Beijing University of Posts and Telecommunications)