Buy the Rumor, Sell the News: When Is News Priced In?
arXiv:2608. 14014v1 Announce Type: new Abstract: Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold.
arXiv:2608. 14014v1 Announce Type: new Abstract: Two old market sayings hold that news is already priced in by the time it is published, and that the rumor is bought while the news is sold.
arXiv:2607. 20645v1 Announce Type: cross Abstract: We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements.
arXiv:2607. 20441v1 Announce Type: cross Abstract: Every information ecosystem produces beliefs that shape strategic decisions.
The paper audits the impact of temporal leakage on financial-news direction prediction across 49,799 articles and 16 feature-model combinations, including TF‑IDF, MiniLM, FinBERT, and fine‑tuned RoBERTa‑large / DeBERTa‑v3‑large, as well as zero/few‑shot and LoRA probes of Llama‑3 and Qwen2.5. Random train‑test splits inflate MCC scores by 1.1× to 6.5×, with larger models and richer features showing greater gains, while end‑to‑end FinBERT fine‑tuning actually increases the gap. Only the mergers and acquisitions (M&A) category shows a positive locked‑test signal under near‑temporal chronological evaluation, with the signal localized to 2024‑2025 European‑tilted M&A semantics and not transferring to a 2009‑2020 U.S. corpus.
The paper investigates the coherence of probabilistic forecasts produced by language models, particularly in the context of life‑decision support. Using a de Finetti‑based method, the authors elicit forecasts for events derived from stock return data and compute the maximum Dutch‑book profit via linear programming, which quantifies incoherence. They find significant incoherence, especially when events have complex logical relationships or when irrelevant context is present, and suggest that alternative training strategies could improve coherence.
The paper proposes measuring a language model’s understanding via no‑arbitrage, defining it as the inability of a bounded trader to profit from Dutch books against the model’s probabilities on logically related claims. It shows that full logical coherence is computationally infeasible, that standard next‑token training yields incoherent predictions across formats, and that uncertainty grows predictably along reasoning chains, creating arbitrage opportunities. The authors introduce Arbitr, a training framework that penalizes logical inconsistencies while maintaining accuracy, reducing exploitability by orders of magnitude and revealing a scaling illusion where large models appear coherent yet exhibit extreme unjustified confidence.
The paper investigates the coherence of probabilistic forecasts produced by language models, particularly when users rely on them for life decisions involving uncertain events. Using a de Finetti-based method, the authors extract forecasts from language models about stock‑return events and compute the maximum Dutch‑book profit via linear programming, a metric of incoherence that does not require observed outcomes. The study finds significant incoherence, especially when events have complex logical relationships or when irrelevant context is present, and suggests that alternative training strategies could improve probabilistic coherence.
arXiv:2608.22852v1 Announce Type: new Abstract: Large language models (LLMs) are increasingly used in investment decision-making, yet prior work shows that they exhibit systematic, model-specific inv...
Alpha‑R1 introduces a reinforcement‑learning aligned large language model framework that performs context‑aware alpha screening by semantically gating candidate factors against a dynamic market state description. The model, trained with group relative policy optimization using realized portfolio returns as reward, selects a sparse subset of factors whose economic rationale matches current market conditions. In a 12‑month out‑of‑sample test, Alpha‑R1 achieved annualized returns of 47.87% on the S&P 500 and 40.57% on the CSI 300, with Sharpe ratios of 1.62 and 2.23, demonstrating the effectiveness of semantic factor reranking in non‑stationary markets.
arXiv:2607. 20449v1 Announce Type: cross Abstract: LLMs are trained predominantly on human-authored text, yet the structural and narrative conventions embedded in that text are rarely examined as a source of systematic behavioral influence, or as a governance risk in deployed systems.
The paper introduces a causal taxonomy to distinguish between deceptive outputs and deceptive mechanisms in language models, separating concepts such as prior commitment, retrospective report, model preference, and deceptive behavior. Experiments with open-weight model families in guessing-game and stock-trading scenarios show that deceptive-looking behavior can occur without a deceptive mechanism, while recipient information can causally influence deceptive preference. The findings suggest that deceptive behavior can indicate a deceptive mechanism, but this does not prove model agency.
arXiv:2609.23703v1 Announce Type: cross Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...