arXiv AI By Hamid Boustanifar, Sasan Mansouri

Same Text, Different Numbers: The Divergence of LLM-Based Measures

Read the original on arXiv AI →

Researchers investigated how different large language models (LLMs) convert corporate text into empirical variables, focusing on thirteen measures such as sentiment, management clarity, uncertainty, answer specificity, and climate and political risk. Using seven LLMs to score earnings call transcripts of S&P 500 companies, they found low cross-model rank correlations (average 0.52) and that transcript-level differences across providers explained only 34% of total score variation. The study shows that model choice significantly alters downstream inference, with varying coefficient magnitudes, signs, and statistical significance, and that averaging across providers stabilizes rankings but not score levels.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv Computation and Language
Aug 27

Anchoring Bias in LLM-as-a-Judge Systems: Prior Scores Compromise Evaluation Independence

The study investigates how prior scores influence large language model (LLM) judgments in the LLM-as-a-Judge paradigm. By testing three prompt conditions—no metadata, revision framing, and anchored metadata containing prior scores—the authors find that prior scores systematically bias evaluations, shifting ratings toward those scores across 192,000 attempts. The bias also affects categorical decisions, blocking 48% of error corrections and flipping 10.18% of correct judgments, and is not mitigated by Chain-of-Thought or a warning, underscoring the need for careful context engineering.

By Ante Kapetanovic, Kemal Altwlkany, Andro Mercep, Tomislav Duricic, Emanuel Lacic
arXiv Machine Learning
Sep 14

Aligning Language Models with Observational Data: Opportunities and Risks from a Causal Perspective

The paper discusses how large language models (LLMs) can be fine‑tuned with observational data to improve alignment with human preferences and business goals. It highlights that directly using such data can cause models to learn spurious correlations, and introduces DeconfoundLM, a method that removes known confounders from reward signals. Experiments show that DeconfoundLM better recovers causal relationships and outperforms baseline methods by over 16% in objective score when confounding is present.

By Erfan Loghmani
arXiv Computation and Language
Sep 11

Same Day, Same Story; One Day Ahead, a Different Signal: The Dual Validity of Financial Sentiment

The study examines whether financial sentiment tools that are validated against human labels also reliably predict market outcomes. Using a large corpus of securities class action messages linked to abnormal stock returns, the authors compare five sentiment instruments—VADER, Loughran‑McDonald, FinBERT, Twitter‑RoBERTa, and an LLM annotator—within a single pipeline. Results show that the alignment between human agreement and sentiment scores varies with sampling strategy and time horizon: conventional sampling favors same‑day associations, while fixed‑n panels yield similar correlations for both same‑day and one‑day‑ahead predictions, yet overall predictive rankings remain weak.

By AS Aravinthkakshan, Laven Srivastava, Harsh Nandwani
arXiv AI
2d ago

A Citation-Grounded Benchmark for Trustworthy Earnings Call Transcript Analysis with Large Language Models

The paper introduces a benchmark for evaluating large language models on trustworthy analysis of earnings call transcripts. It proposes a numeric evidence evaluation method that assesses groundedness without expert annotation, and presents an automated pipeline that builds the ECTs-100 dataset from the top 100 S&P 500 constituents. The study also explores the failure mode of conscious incompetence, where models must recognize insufficient evidence and avoid hallucinations, finding that while groundedness is strong, correctness remains a challenge.

By Yingzhu Zhao, Vlad Pandelea, Han Yuan, Bo Hu, Wuqiong Luo, Li Zhang, Zheng Ma