The study examines whether large language models (LLMs) can accurately simulate individual financial users by conducting a longitudinal paper‑trading experiment with 80 participants. Using a rolling next‑day prediction protocol, the researchers compared LLM predictions to a simple recent‑activity persistence baseline across multiple behavioral fidelity levels, from trade occurrence to asset selection and portfolio outcomes. Results show that no LLM consistently outperforms the baseline, with fidelity decreasing at finer behavioral granularity, and that recent trading history largely drives activity predictions while asset selection depends more on available evidence.
By Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu
The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.
By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
arXiv:2606. 08285v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling.
By Junyi Yao, Zihao Zheng
arXiv:2601.15322v3 Announce Type: replace-cross
Abstract: Tool-using agents can repeat a final decision while changing their recorded execution. We introduce the Determinism-Faithfulness Assurance Ha...
By Raffi Khatchadourian
arXiv:2607. 20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time.
By Raffi Khatchadourian
arXiv:2608.29372v1 Announce Type: new
Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a singl...
By Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie