arXiv Computation and Language
1d ago

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

The study evaluates how well three sentiment‑analysis tools (TextBlob, VADER, Twitter‑roBERTa‑base) and three large language models (Qwen3‑32B, GPT‑OSS‑120B, Llama‑4‑Maverick‑17B) agree with six human raters on 100 tweets. Agreement was measured with Cohen’s and Fleiss’ kappa, revealing only fair inter‑human agreement and higher concordance for binary sentiment labels than for three‑class labels. Twitter‑roBERTa‑base achieved the strongest alignment with humans, especially for negative versus non‑negative sentiment, while the LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, particularly for positive versus non‑positive classifications. The findings emphasize that domain‑specific fine‑tuning and human‑centered evaluation are essential for reliable social media sentiment analysis.

By Himarsha R. Jayanetti, Sivakanesan Dhanushkanda, Shuai Hao, Michael L. Nelson, Michele C. Weigle
arXiv AI
Aug 7

Beyond Sentiment: Comparing Traditional NLP and LLM-Based Multi-Dimensional Analysis for Political News Evaluation

arXiv:2608. 05155v1 Announce Type: cross Abstract: Traditional sentiment analysis (SA) models, while effective for polarity classification, provide limited insight into the rhetorical, ideological, and framing dimensions of political discourse -- dimensions that are central to research in the social sciences and humanities (SSH).

By Maryam Fooladi, Federico Bottino
OpenAI Blog
Apr 6, 2017

Unsupervised sentiment neuron

We’ve developed an unsupervised system which learns an excellent representation of sentiment, despite being trained only to predict the next character in the text of Amazon reviews.