arXiv Computation and Language By Himarsha R. Jayanetti, Sivakanesan Dhanushkanda, Shuai Hao, Michael L. Nelson, Michele C. Weigle

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

Read the original on arXiv Computation and Language →

The study evaluates how well three sentiment‑analysis tools (TextBlob, VADER, Twitter‑roBERTa‑base) and three large language models (Qwen3‑32B, GPT‑OSS‑120B, Llama‑4‑Maverick‑17B) agree with six human raters on 100 tweets. Agreement was measured with Cohen’s and Fleiss’ kappa, revealing only fair inter‑human agreement and higher concordance for binary sentiment labels than for three‑class labels. Twitter‑roBERTa‑base achieved the strongest alignment with humans, especially for negative versus non‑negative sentiment, while the LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, particularly for positive versus non‑positive classifications. The findings emphasize that domain‑specific fine‑tuning and human‑centered evaluation are essential for reliable social media sentiment analysis.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Computation and Language
Sep 11

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
arXiv AI
Aug 5

How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.

By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv AI
Aug 7

Automatic Detection of Deaths from Social Networking Sites

arXiv:2608. 05183v1 Announce Type: cross Abstract: This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media content, and reported machine learning (ML) classifiers that achieved high performance in automatically detecting deaths of social networking site users from posts associated with their profiles.

By Nuhu Ibrahim, Riza Batista-Navarro