arXiv AI

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.

arXiv Computation and Language
4d ago

Can Language Models Learn to Forecast Stock Prices

arXiv:2609.36914v1 Announce Type: new Abstract: Post-training has been shown to significantly improve language models' performance on tasks with verifiable outcomes, including mathematical reasoning,...

By Jiacheng Guo, Suozhi Huang, Shuzhen Li, Yunlong Gao, Zerui Cheng, Jason Ge, Shushu Liang, Zihao Li, Hao Lu, Ming Yin, Shilong Liu, Jiashuo Liu, Xu Kuang, Mengdi Wang
arXiv AI
Aug 28

The Reasoning Tax: Token Economics of LLM Reasoning Across Task Types and Deployment Contexts

The paper introduces the Token Economy Score (TES), a metric that quantifies the accuracy gain of reasoning-capable large language models relative to non-reasoning baselines, normalized by token generation cost. An empirical study across 151 runs on seven diverse benchmarks shows that task structure—such as sequential inference chains—predicts higher TES, while knowledge-recall tasks yield lower TES despite difficulty. The analysis also reveals diminishing returns at higher reasoning effort and highlights how deployment context, via Reasoning Cost Share and Deployment Cost Multiplier, can alter the economic viability of reasoning workloads.

By Sachin Gopal Wani, Ajay Dholakia, David Ellison
arXiv Machine Learning
Sep 25

A Fast and Effective Solution to the Problem of Look-ahead Bias in LLMs

The paper addresses look‑ahead bias in large language models (LLMs) used for financial prediction, which arises because LLMs are trained on long time‑series data. It proposes a low‑cost solution that adjusts the logits of a base model at inference time using two smaller, specialized models—one fine‑tuned to forget certain information and another to retain it. Experiments show that this method removes both verbatim and semantic knowledge, corrects biases, and outperforms previous approaches.

By Humzah Merchant, Bradford Levy
arXiv AI
Jun 4

FinTradeBench: A Financial Reasoning Benchmark for LLMs

arXiv:2603. 19225v3 Announce Type: replace-cross Abstract: Real-world financial decision-making is a challenging problem that requires reasoning over heterogeneous signals, including company fundamentals derived from regulatory filings and trading signals computed from price dynamics.

By Yogesh Agrawal, Aniruddha Dutta, Md Mahadi Hasan, Santu Karmaker, Aritra Dutta
arXiv AI
Jul 14

Can Agentic Trading Systems Pay for Their Own Intelligence?

arXiv:2607. 10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value.

By Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
arXiv AI
3d ago

Reasoning Externalization for Faithful Large Language Model Narratives of Stock Return Predictions

The paper introduces a framework that uses large language models (LLMs) to generate natural‑language narratives explaining cross‑sectional stock return predictions. It combines temporal Shapley additive explanations (SHAP) from an XGBoost model with historical regime analogs to provide context. A controlled study shows that progressively externalizing numerical and relational reasoning improves evidence faithfulness and accuracy, while historical analogs boost human‑rated usefulness.

By Sujung Kim, Seung Hwan Cho, Sangjin Park, Young-Min Kim