arXiv Machine Learning

From Financial Sentiment Classification to Return Predictability: A QLoRA Benchmark of Large Language Models

arXiv:2608. 04200v1 Announce Type: cross Abstract: Financial sentiment classifiers are commonly evaluated against human labels, but strong linguistic performance does not necessarily imply economically useful return predictability.

arXiv Machine Learning
Sep 18

Evaluating Financial Sentiment in the Age of AI

The paper evaluates twelve financial sentiment models—including dictionary-based methods, finance-specific transformers, and open-source large language models—using linguistic and economic validity criteria. General-purpose LLMs match finance-specific transformers in classification performance but do not yield stronger economic relationships. While several models correlate with earnings surprises, none shows a significant link to next‑day stock returns, and performance is strongest for large earnings beats or misses.

By Arslan Bisharat, Oudom Hean
arXiv Computation and Language
Sep 11

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
arXiv Computation and Language
Aug 21

Reliable Financial Named Entity Recognition under Domain Shift

arXiv:2608. 19558v1 Announce Type: new Abstract: Financial AI systems often train information extractors on one textual register and deploy them across filings, news, and user-generated content, while standard F1 scores do not indicate which predictions remain safe to automate when the input distribution changes.

By Zihao Zheng, Baichuan Li, Junyi Yao, Jiayu Long
arXiv Computation and Language
Aug 27

Reliable Financial Named Entity Recognition Under Domain Shift: Confidence Estimation and Selective Prediction

The paper investigates confidence estimation and selective prediction for financial named entity recognition (NER) under domain shift, using a stress test across SEC filings, financial news, and social media. It evaluates BERT and LoRA‑tuned Qwen2.5 models with five inference‑time confidence signals, finding that whole‑output probability is a strong in‑domain error detector but weak out‑of‑domain, while entity‑span probability and self‑consistency remain robust. Abstention can dramatically reduce sentence error on high‑confidence in‑domain data, but offers limited benefit under extreme social‑media shift, suggesting a staged deployment that first detects severe distribution shift before applying confidence gating.

By Zihao Zheng, Baichuan Li, Junyi Yao, Jiayu Long
Hugging Face Trending Papers
Jul 27

LLM-Based vs. Lexicon-Based Sentiment Signals for Tail-Risk Detection in Meme Stocks

This paper presents an empirical comparison of lexicon-based and Large Language Model (LLM)-based sentiment analysis for extracting market-relevant signals from social media discourse in highly volatile equity markets. Using Reddit data from r/WallStreetBets and focusing on meme stocks (GME, AMC, NOK), we construct time-aligned sentiment indicators and evaluate their relationship with market returns, with particular attention to extreme positive return events in the upper tail of the return distribution.