arXiv Computation and Language

Nobody Truly Agrees on Sentiment: Humans, Bespoke Tools, and LLMs Struggle with Social Media Texts

The study evaluates how well three sentiment‑analysis tools (TextBlob, VADER, Twitter‑roBERTa‑base) and three large language models (Qwen3‑32B, GPT‑OSS‑120B, Llama‑4‑Maverick‑17B) agree with six human raters on 100 tweets. Agreement was measured with Cohen’s and Fleiss’ kappa, revealing only fair inter‑human agreement and higher concordance for binary sentiment labels than for three‑class labels. Twitter‑roBERTa‑base achieved the strongest alignment with humans, especially for negative versus non‑negative sentiment, while the LLMs showed substantial agreement among themselves and moderate to substantial alignment with humans, particularly for positive versus non‑positive classifications. The findings emphasize that domain‑specific fine‑tuning and human‑centered evaluation are essential for reliable social media sentiment analysis.

arXiv Computation and Language
Sep 11

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
arXiv AI
Aug 5

How Closely Do LLM Reviews Align with Human Peer Review?

arXiv:2608. 03659v1 Announce Type: cross Abstract: Large language models (LLMs) are increasingly used to generate scientific reviews, yet existing evaluations rarely examine whether different providers align with both conference decisions and human reviewing priorities within the same controlled setting.

By Abraham Camelo-Guerrero, Jairo Diaz-Rodriguez
arXiv AI
Aug 7

Automatic Detection of Deaths from Social Networking Sites

arXiv:2608. 05183v1 Announce Type: cross Abstract: This dissertation analysed and discussed the differences in linguistic characteristics between pre-mortem and post-mortem social media content, and reported machine learning (ML) classifiers that achieved high performance in automatically detecting deaths of social networking site users from posts associated with their profiles.

By Nuhu Ibrahim, Riza Batista-Navarro
arXiv AI
Sep 25

Human Agreement and Return Association Are Not Interchangeable Criteria

The paper examines whether human agreement and return association can be used interchangeably as criteria for validating sentiment tools in financial NLP. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments and find that the relationship between human agreement and predictive validity varies with sampling conventions and score representations. They conclude that benchmark agreement establishes semantic validity but does not guarantee predictive rankings, and that message volume in a spam‑heavy conversation does not predict market damage or settlement size.

By AS Aravinthakshan, Laven Srivastava, Harsh Nandwani
arXiv AI
Aug 5

VIBE: A VAD-Informed Benchmark for Entity-Centered Affective Profiling of Large Language Model Outputs

arXiv:2608. 03810v1 Announce Type: cross Abstract: Large language models routinely describe socially salient targets, including political figures, countries, religions, organizations, historical events, and social groups, encoding affective framing alongside factual content: a target may appear favorable or threatening, calm or conflictual, powerful or vulnerable.

By Andrei Chetvergov, Alexander Evseev, Timofei Sivoraksha, Stepan Ukolov, Mikhail Solovev, Danil Sazanakov, Sergey Bolovtsov