arXiv Computation and Language

ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

ViTOED is a new dataset for target‑oriented emotion detection in Vietnamese social media, containing 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity). The dataset uncovers Vietnamese‑specific linguistic phenomena such as implicit sources and targets and vocabulary ambiguities, and it serves as a benchmark for evaluating Vietnamese pre‑trained language models. A baseline using structured sentiment graphs shows that span detection and relation extraction remain challenging, indicating significant room for improvement in Vietnamese target‑oriented emotion detection tasks.

arXiv Computation and Language
Aug 27

KESA: A Knowledge Enhanced Approach For Sentiment Analysis

The paper introduces KESA, a knowledge‑enhanced approach for sentence‑level sentiment analysis that incorporates sentiment knowledge through two auxiliary tasks: sentiment word cloze and conditional sentiment prediction. These tasks use prior sentiment polarity to guide the selection of sentiment words and the prediction of overall sentiment, respectively, and explore label combination methods to unify multiple label types. Experiments show that KESA consistently outperforms pre‑trained models and complements existing knowledge‑enhanced post‑training methods.

By Qinghua Zhao, Shuai Ma, Shuo Ren
arXiv AI
3d ago

BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects

The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.

By Ali Almutairi, Gelareh Mohammadi, Imran Razzak, Aditya Joshi