arXiv Computation and Language By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

Read the original on arXiv Computation and Language →

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv AI
Sep 25

Human Agreement and Return Association Are Not Interchangeable Criteria

The paper examines whether human agreement and return association can be used interchangeably as criteria for validating sentiment tools in financial NLP. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments and find that the relationship between human agreement and predictive validity varies with sampling conventions and score representations. They conclude that benchmark agreement establishes semantic validity but does not guarantee predictive rankings, and that message volume in a spam‑heavy conversation does not predict market damage or settlement size.

By AS Aravinthakshan, Laven Srivastava, Harsh Nandwani
arXiv Machine Learning
Sep 18

Evaluating Financial Sentiment in the Age of AI

The paper evaluates twelve financial sentiment models—including dictionary-based methods, finance-specific transformers, and open-source large language models—using linguistic and economic validity criteria. General-purpose LLMs match finance-specific transformers in classification performance but do not yield stronger economic relationships. While several models correlate with earnings surprises, none shows a significant link to next‑day stock returns, and performance is strongest for large earnings beats or misses.

By Arslan Bisharat, Oudom Hean