Hugging Face Blog

SetFitABSA: Few-Shot Aspect Based Sentiment Analysis using SetFit

arXiv AI
Oct 1

BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects

The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.

By Ali Almutairi, Gelareh Mohammadi, Imran Razzak, Aditya Joshi
arXiv Computation and Language
4d ago

Mawqif-XT: An Arabic Benchmark Dataset for Cross-Target Stance Detection

The paper introduces Mawqif-XT, a new Arabic benchmark dataset comprising 996 manually annotated tweets from three public targets: Women Driving, E-Cars, and Trimester System. Each tweet is labeled for stance, sentiment, and sarcasm following the Mawqif annotation scheme, and the dataset is intended as a held‑out evaluation set to test cross‑target generalization. Baseline results are provided using Arabic and multilingual transformer models as well as zero‑shot large language models, enabling reproducible evaluation alongside the original Mawqif dataset.

By Rasha Albalawi, Nuha Albadi, Hamzah Luqman, Maram Kurdi, Saad Ezzini, Asma Yamani, Ahmed Ashraf
arXiv Computation and Language
Sep 18

ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

ViTOED is a new dataset for target‑oriented emotion detection in Vietnamese social media, containing 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity). The dataset uncovers Vietnamese‑specific linguistic phenomena such as implicit sources and targets and vocabulary ambiguities, and it serves as a benchmark for evaluating Vietnamese pre‑trained language models. A baseline using structured sentiment graphs shows that span detection and relation extraction remain challenging, indicating significant room for improvement in Vietnamese target‑oriented emotion detection tasks.

By Chanh Vo, Son T. Luu, Ngan Luu-Thuy Nguyen
arXiv AI
Sep 3

DeepAffinity: Long-Term Aspect Preference Prediction in eCommerce using Small Language Models

DeepAffinity is a model designed to predict eCommerce users’ future preferences for product aspects such as brand, size, and color, treating this as a temporal prediction problem. It uses small language models with structured prompts and specialized prediction heads, outperforming standard generative fine‑tuning and general‑purpose open‑source LLMs that lack task‑specific tuning. The approach improves recommendation quality on a large multinational eCommerce platform.

By Yotam Eshel, Guy Hadad, Guy Feigenblat, Yuri M. Brovman, Matt Gearhart, Bracha Shapira