arXiv Computation and Language By Ping Kuen Wong

DisclosureBeta: A Measurement-Channel Theory for Regime-Conditioned Betas from LLM-Read Risk Disclosures

Read the original on arXiv Computation and Language →

The paper introduces DisclosureBeta, a theory that treats large language models (LLMs) as noisy measurement channels for a firm’s latent risk characteristics, integrating this noise into the asset‑pricing error budget. It establishes identification and consistency of regime‑conditional beta loadings within a piecewise‑stationary Fama‑French five‑factor framework, provides matching lower bounds, and proposes an adaptive estimator that blends text‑based and rolling‑window approaches, improving precision when price histories are short or regime‑breaks occur. The work also outlines a pre‑registered empirical evaluation on firms with thin price histories.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv Computation and Language.

arXiv Machine Learning
Aug 19

Temporal Leakage in Financial News NLP: A Multi-Architecture Audit with a Regime-Specific M&A Signal

The paper audits the impact of temporal leakage on financial-news direction prediction across 49,799 articles and 16 feature-model combinations, including TF‑IDF, MiniLM, FinBERT, and fine‑tuned RoBERTa‑large / DeBERTa‑v3‑large, as well as zero/few‑shot and LoRA probes of Llama‑3 and Qwen2.5. Random train‑test splits inflate MCC scores by 1.1× to 6.5×, with larger models and richer features showing greater gains, while end‑to‑end FinBERT fine‑tuning actually increases the gap. Only the mergers and acquisitions (M&A) category shows a positive locked‑test signal under near‑temporal chronological evaluation, with the signal localized to 2024‑2025 European‑tilted M&A semantics and not transferring to a 2009‑2020 U.S. corpus.

By Chenhao Xue, Raslen Guesmi, Siwei Feng, Yucheng Gong, Jacob Xavier Sundram, Jordan Pang, Lan Wang, Julian Kaljuvee