Large language models (LLMs) are being tested as simulators of individual financial decision-making. In a controlled paper‑trading experiment with 120 volunteers, the study evaluated whether an LLM could predict a participant’s next‑day trading action, chosen security, and transaction size using only pre‑cutoff information. Results showed that including market context improved predictions of actions and tickers, but sizing remained challenging, and the models tended to over‑predict hold actions, under‑predict sells, and simplify multi‑security trades.
By Jiajie He, Jiangyuan Hong, Dongling Ni, Wenjin Liu, Xintong Chen
arXiv:2605. 28850v2 Announce Type: replace Abstract: We study behavioral alignment and representation dynamics of large language model (LLM) agents in financial decision environments.
By Weicheng Xue
arXiv:2606. 02798v1 Announce Type: new Abstract: Many decision-support settings require systems that adapt to individual users, but evaluation data for this problem remain limited.
By Liangwei Yang, Jielin Qiu, Zixiang Chen, Ming Zhu, Juntao Tan, Zhiwei Liu, Wenting Zhao, Zhujun Lan, Akshara Prabhakar, Silvio Savarese, Huan Wang, Shelby Heinecke
arXiv:2608. 04095v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used as personalized assistants in high-stakes domains such as financial advising, yet it remains unclear whether they can maintain and update an individualized user model over long horizons.
By Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
arXiv:2606. 31461v1 Announce Type: new Abstract: Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules.
By Yao Shi, Kingfung Luo, Nan Tang, Yuyu Luo
The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.
By Jiayi Chen, Guiling Wang
arXiv:2407. 18957v5 Announce Type: replace-cross Abstract: Can AI Agents simulate real-world trading environments to investigate the impact of external factors on stock trading activities (e.
By Chong Zhang, Xinyi Liu, Zhongmou Zhang, Mingyu Jin, Lingyao Li, Zhenting Wang, Wenyue Hua, Dong Shu, Suiyuan Zhu, Xiaobo Jin, Sujian Li, Mengnan Du, Yongfeng Zhang
arXiv:2608.29372v1 Announce Type: new
Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a singl...
By Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie
arXiv:2502. 18834v3 Announce Type: replace-cross Abstract: Financial time series (FinTS) record the behavior of human-brain-augmented decision-making, capturing valuable historical information that can be leveraged for profitable investment strategies.
By Yifan Hu, Yuante Li, Peiyuan Liu, Yuxia Zhu, Naiqi Li, Tao Dai, Shu-tao Xia, Dawei Cheng, Changjun Jiang
EvolveTrade is a self‑evolving framework that treats the system prompt of a tool‑using LLM trading agent as a text‑parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and portfolio feedback while keeping the backbone LLM fixed, allowing the agent to refine its information‑acquisition and portfolio‑construction procedures over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed‑policy baselines, with behavioral analyses indicating increased code‑mediated analysis and regime‑relevant computations.
By Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang
Large language models (LLMs) can synthesize financial narratives but may express high confidence when evidence is sparse, stale, or contradictory. This failure is especially consequential in forecasting, where filings, news, prices, volume, and technical signals can disagree.
arXiv:2606. 31522v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment.
By Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie