arXiv Machine Learning

SalAngaBhava: A Sinhala Market Dataset for Aspect-based Sentiment Analysis

arXiv:2607. 05259v1 Announce Type: cross Abstract: Sentiment analysis has been a primary domain under Natural Language Processing (NLP) from its inception as it plays a vital role in both real-world and research applications.

arXiv AI
3d ago

BARRAC: Adaptation of an English Aspect-based Sentiment Analysis Approach for Classification Tasks in Arabic Dialects

The paper introduces BARRAC, a method that adapts an English aspect‑based sentiment analysis framework for Arabic dialect classification tasks. It replaces English consumer‑review attribute pools with Arabic linguistic markers for sentiment, sarcasm, and dialect identification, and swaps noisy self‑training for a two‑stage training process. Evaluated on five Arabic dialect datasets, BARRAC achieves a mean macro‑F1 of 63.93%, surpassing the best few‑label state‑of‑the‑art by 3% and outperforming GPT‑4o on four of the five tasks, while error analysis highlights remaining challenges.

By Ali Almutairi, Gelareh Mohammadi, Imran Razzak, Aditya Joshi
arXiv Computation and Language
Sep 18

ViTOED: A Dataset for Target-Oriented Emotion Detection on Vietnamese Social Media Texts

ViTOED is a new dataset for target‑oriented emotion detection in Vietnamese social media, containing 10,985 user comments and 21,244 manually annotated opinion quadruples (source, target, expression, polarity). The dataset uncovers Vietnamese‑specific linguistic phenomena such as implicit sources and targets and vocabulary ambiguities, and it serves as a benchmark for evaluating Vietnamese pre‑trained language models. A baseline using structured sentiment graphs shows that span detection and relation extraction remain challenging, indicating significant room for improvement in Vietnamese target‑oriented emotion detection tasks.

By Chanh Vo, Son T. Luu, Ngan Luu-Thuy Nguyen