arXiv Computation and Language

Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection

The paper introduces Mawqif-XT, a new Arabic benchmark dataset comprising 996 manually annotated tweets from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is labeled for stance, sentiment, and sarcasm following the Mawqif annotation scheme, and the dataset is intended as a held‑out evaluation set to test cross‑target generalization. Baseline results are provided using Arabic and multilingual transformer models as well as zero‑shot large language models, enabling reproducible evaluation alongside the original Mawqif dataset.

arXiv Machine Learning
Sep 25

TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)

The paper introduces CLASP‑Ar, a cloze‑style prompting method for Arabic stance detection that replaces complex multitask learning and ensembles with a single masked language modeling prompt. By combining the target, predicted sentiment, and text into one prompt and constraining the [MASK] prediction to a verbalizer‑defined label set, the approach aims to simplify the task while maintaining performance.

By Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler
arXiv AI
Oct 1

BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects

The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.

By Ali Almutairi, Gelareh Mohammadi, Imran Razzak, Aditya Joshi
arXiv Computation and Language
3d ago

StanceEval 2026: The Second Stance Detection Shared Task

StanceEval 2026 is the second edition of a shared task on stance detection in Arabic social media, where systems must classify a tweet’s stance toward a target as Favor, Against, or None. The event featured two tracks: Track 1 tests cross‑target transfer on thematically related topics (Women Driving vs. Women Empowerment), while Track 2 evaluates cross‑domain transfer to entirely unseen targets (E‑Cars and Trimester System). With 80 registered teams and 30 submissions, top systems achieved $F_{avg2}$ scores of 0.8994 (Track 1) and 0.9400 (Track 2), surpassing baseline performance and highlighting challenges such as target polarization and dialectal nuance.

By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif
Hugging Face Trending Papers
Jul 29

AHA-Memes: A Fine-Grained Multimodal Benchmark for Understanding Hate in Arabic Memes

Hateful memes are a growing form of multimodal online harm, where hostile intent is often conveyed through the joint interpretation of images, text, cultural references, and implicit targets. While hateful meme detection has advanced in high-resource languages, Arabic remains underexplored, with existing meme resources focusing mainly on propaganda or coarse harmful-content labels.

arXiv Machine Learning
Sep 11

E-CONAN (Entailment, CONtradition And Neutral) Benchmarks: Arabic Textual Entailment and Natural Inference Datasets

E-CONAN introduces Arabic textual entailment and natural inference benchmarks comprising two datasets: E-CONAN-2 (2-way RTE) and E-CONAN-3 (3-way NLI). The datasets are built from automatically-translated pairs, human-validated machine translations, hand-crafted pairs from Arabic teaching books, and rumor-containing news headlines. The authors evaluated nine multilingual pretrained models and five large language models on these benchmarks, demonstrating that E-CONAN offers a more diverse and robust assessment than existing datasets like XNLI and ArNLI.

By Khloud AL Jallad, Nada Ghneim, Ghaida Rebdawi
arXiv AI
Oct 1

TTLab at Daleel 2026: STAR-Ar, Sequence Tagging for Argument Recognition in Arabic

The paper introduces STAR‑Ar, a BERT‑BiLSTM‑CRF model designed for the Daleel 2026 Arabic argument mining shared task. It treats argument discourse unit detection and classification as a token‑level sequence labeling problem, achieving an F1‑score of 72.69 on validation and 73.7 on test data. Analysis shows that models trained only on editorial texts perform worse than those trained on debates, mainly due to the smaller editorial dataset.

By Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler