arXiv Computation and Language By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Asma Yamani, Maram Kurdi, Saad Ezzini, Ahmed Ashraf, Maged Al-Shaibani, Nora Alturayeif

StanceEval 2026: The Second Stance Detection Shared Task

Read the original on arXiv Computation and Language →

StanceEval 2026 is the second edition of a shared task on stance detection in Arabic social media, where systems must classify a tweet’s stance toward a target as Favor, Against, or None. The event featured two tracks: Track 1 tests cross‑target transfer on thematically related topics (Women Driving vs. Women Empowerment), while Track 2 evaluates cross‑domain transfer to entirely unseen targets (E‑Cars and Trimester System). With 80 registered teams and 30 submissions, top systems achieved $F_{avg2}$ scores of 0.8994 (Track 1) and 0.9400 (Track 2), surpassing baseline performance and highlighting challenges such as target polarization and dialectal nuance.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
3d ago

Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection

The paper introduces Mawqif-XT, a new Arabic benchmark dataset comprising 996 manually annotated tweets from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is labeled for stance, sentiment, and sarcasm following the Mawqif annotation scheme, and the dataset is intended as a held‑out evaluation set to test cross‑target generalization. Baseline results are provided using Arabic and multilingual transformer models as well as zero‑shot large language models, enabling reproducible evaluation alongside the original Mawqif dataset.

By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Maram Kurdi, Saad Ezzini, Asma Yamani, Ahmed Ashraf
arXiv AI
Aug 5

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.

By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov
arXiv Machine Learning
Sep 25

TTLab at StanceEval-2026: A Cloze-Style Prompting Approach for Arabic-Language Stance Detection (CLASP-Ar)

The paper introduces CLASP‑Ar, a cloze‑style prompting method for Arabic stance detection that replaces complex multitask learning and ensembles with a single masked language modeling prompt. By combining the target, predicted sentiment, and text into one prompt and constraining the [MASK] prediction to a verbalizer‑defined label set, the approach aims to simplify the task while maintaining performance.

By Bhuvanesh Verma, Ali Abusaleh, Alexander Mehler
arXiv AI
Sep 7

MABPD: Multi-Agent Bias Probing & Detection via Structured Argument Debate

MABPD (Multi‑Agent Bias Probing & Detection) is a training‑free pipeline that uses three specialized large language model agents to analyze news articles from complementary perspectives and resolve disagreements via a Structured Argument Debate (SAD) protocol. SAD imposes an asymmetric burden of proof—biased claims lacking grounded textual evidence receive zero weight—along with role‑weighted voting and post‑consensus verification, replacing task‑specific supervised decision boundaries. Ablation studies show that the debate module alone accounts for up to a 10.6‑point F1 gain, and on the BABE benchmark MABPD attains 83.4% macro F1, within 0.7 percentage points of the supervised state‑of‑the‑art, while achieving 75.0% zero‑shot accuracy on the SemEval 2019 HyperPartisan corpus.

By Garvit Joshi (Graphic Era University, Dehradun, India), Stavya Dhyani (Graphic Era University, Dehradun, India), Jasmine (Graphic Era University, Dehradun, India), Arun Chauhan (Graphic Era University, Dehradun, India)
arXiv Computation and Language
Oct 1

Halluscoring 2026: The first shared task on llms hallucination detection and answer verification

HalluScoring 2026 is a shared task that evaluates hallucination detection and factual verification in Arabic question answering, focusing on generalization to unseen questions and LLMs. It comprises two main tasks with four subtasks: binary hallucination detection (Subtasks 1.1 and 1.2) and answer verification against six candidates in Islamic and general knowledge domains (Subtasks 2.1 and 2.2). Thirteen teams participated, with the top system achieving AUC‑ROC scores of 0.772 and 0.767 for detection, and 0.882 and 0.857 for verification.

By Aisha Alansari, Abdessalam Bouchekif, Ahmed Hasanaath, Salah Eddine Bekhouche, Malak Alkhorasani, Mohammed-En-Nadhir Zighem, Saad Ezzini, Hichem Telli, Hend Al-Khalifa, Muhammad Abdul-Mageed, Hadid Abdenour, Hamzah Luqman