arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
By Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang
arXiv:2609.30316v1 Announce Type: cross
Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already obse...
By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
arXiv:2605. 30363v2 Announce Type: replace-cross Abstract: Regime shifts in financial markets reorganise the joint dynamics of asset prices and macro variables, breaking any single-regime calibration.
By Mingxuan Yi, Vidal Mehra, Jing Chen, John Cartlidge
arXiv:2609.23703v1 Announce Type: cross
Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...
By Kemal Kirtac
arXiv:2606. 18192v1 Announce Type: new Abstract: As high-quality public web corpora become increasingly exhausted, clean long-context documents have become a scarce and expensive source of training data for large language models (LLMs).
By Nick Bettencourt, Xiaowei Ding, Kay Giesecke
MemGuard-Alpha evaluates whether membership inference attacks (MIA) can detect memorization in large language models (LLMs) used for financial alpha signals. The study combines five MIA methods with a temporal proximity feature and a cross-model disagreement metric, then audits them across seven LLMs, 50 S&P 100 stocks, and 299,600 prompt-model pairs. Findings show that temporal proximity alone perfectly predicts in-sample status, MIA discriminative power largely stems from model scale differences, and filtering based on contamination scores does not improve risk-adjusted performance once transaction costs are considered.
By Anisha Roy, Dip Roy
arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.
By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
The paper audits the widely used ISOT/Kaggle Fake and Real News corpus and finds that extremely high reported accuracies (≈0.98) are largely due to shortcut signals rather than genuine veracity detection. A simple TF‑IDF linear classifier achieves perfect F1 when using only subject metadata, and even after removing metadata, newswire tags, and duplicate documents, the F1 drops only modestly, indicating that editorial style rather than specific tokens drives performance. Under topic‑disjoint and temporal transfer tests, performance collapses, and models transfer poorly to the independent LIAR benchmark, showing that within‑corpus scores reflect source and topic separability, not truth verification.
whyItMatters:"The study demonstrates that current high accuracy metrics on this fake‑news dataset are misleading, highlighting the need for more robust evaluation protocols that guard against shortcut learning."
By Yuvraj Verma
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree.
The paper introduces DisclosureBeta, a theory that treats large language models (LLMs) as noisy measurement channels for a firm’s latent risk characteristics, integrating this noise into the asset‑pricing error budget. It establishes identification and consistency of regime‑conditional beta loadings within a piecewise‑stationary Fama‑French five‑factor framework, provides matching lower bounds, and proposes an adaptive estimator that blends text‑based and rolling‑window approaches, improving precision when price histories are short or regime‑breaks occur. The work also outlines a pre‑registered empirical evaluation on firms with thin price histories.
By Ping Kuen Wong
arXiv:2609.34004v2 Announce Type: replace
Abstract: Equity-relevant news evolves through temporally dependent corporate events, making historical information useful only when event continuity, inform...
By Tong Liu, Lanmiao Liu, Xiang Hu
arXiv:2606. 23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks.
By Mostapha Benhenda