Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts
Read the original on arXiv Computation and Language →The study evaluates how well three sentiment‑analysis tools (TextBlob, VADER, Twitter‑roBERTa‑base) and three large language models (Qwen3‑32B, GPT‑OSS‑120B, Llama‑4‑Maverick‑17B) agree with six human raters on 100 tweets. Agreement was measured with Cohen’s and Fleiss’ kappa, revealing only fair inter‑human agreement and higher concordance for binary sentiment labels than for three‑class labels. Twitter‑roBERTa‑base achieved the strongest alignment with humans, especially for negative versus non‑negative sentiment, while the LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, particularly for positive versus non‑positive classifications. The findings emphasize that domain‑specific fine‑tuning and human‑centered evaluation are essential for reliable social media sentiment analysis.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.