arXiv:2605.24564v2 Announce Type: replace
Abstract: Backtesting large language models (LLMs) on historical financial data is unreliable when their pre-training data include the evaluated events. An L...
By Weixian Waylon Li, Mengyu Wang, Tiejun Ma
arXiv:2512. 23847v2 Announce Type: replace-cross Abstract: We develop a statistical procedure to detect lookahead bias in economic forecasts generated by large language models (LLMs).
By Zhenyu Gao, Wenxi Jiang, Yutong Yan
Specialist training beats generalist scale when forecasting financial statements. To our knowledge, no prior work jointly forecasts complete financial statements beyond one year, yet in a discounted-cash-flow valuation most firm value sits past that window.
arXiv:2510. 03519v3 Announce Type: replace-cross Abstract: Time series reasoning is crucial to decision-making in diverse domains, including finance, energy, and scientific discovery.
By Fangxu Yu, Hongyu Zhao, Tianyi Zhou
arXiv:2609.36914v1 Announce Type: new
Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...
By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv:2609.30316v1 Announce Type: cross
Abstract: Language models used in financial backtests suffer from look-ahead bias, as a model trained on text published after the study period has already obse...
By Seunghan Lee, Jun Seo, Jaehoon Lee, Junhyeok Kang, Sangjun Han, Sungdong Yoo, Minjae Kim, Tae Yoon Lim, Dongwan Kang, Hwanil Choi, Soonyoung Lee, Wonbin Ahn
arXiv:2606. 24950v1 Announce Type: new Abstract: Financial decision-making is contextual: forecasting prices, valuing companies, and assessing event exposure weigh price history, accounting fundamentals, macroeconomic regime, and contemporaneous text.
By Patara Trirat, Jin Myung Kwak, Jay Heo, Heejun Lee, Sung Ju Hwang
The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.
By Jiayi Chen, Guiling Wang
arXiv:2608. 11788v1 Announce Type: cross Abstract: Continual pre-training (CPT) has been widely adopted as a method for domain adaptation in large language models.
By Minjun Kim, Inho Won, Hyeonseok Lim, MinKyu Kim, Junghun Yuk, Wooyoung Go, Jongyoul Park, Jungyeul Park, KyungTae Lim
arXiv:2601.17609v3 Announce Type: replace
Abstract: In domains like medicine and finance, large-scale labeled data is costly and often unavailable, leading to models trained on small datasets that st...
By Sara Rezaeimanesh, Kundan Thind, Farzan Siddiqui, Mohammad M. Ghassemi
arXiv:2607. 11889v1 Announce Type: cross Abstract: Large language models trained on unrestricted internet corpora inevitably embed information from the future, introducing lookahead bias that compromises the validity of backtests and causal inference in finance and the social sciences.
By Bryan Kelly, Semyon Malamud, Johannes Schwab, Teng Andrea Xu
arXiv:2608. 11327v1 Announce Type: new Abstract: Specialist training beats generalist scale when forecasting financial statements.
By Travis L. Johnson, Jiannan Jiang, Soumyabrata Chaudhuri, Yihao Chen, Lauren Falvey, Donal O'Cofaigh