The study examines whether large language models (LLMs) can accurately simulate individual financial users by conducting a longitudinal paper‑trading experiment with 80 participants. Using a rolling next‑day prediction protocol, the researchers compared LLM predictions to a simple recent‑activity persistence baseline across multiple behavioral fidelity levels, from trade occurrence to asset selection and portfolio outcomes. Results show that no LLM consistently outperforms the baseline, with fidelity decreasing at finer behavioral granularity, and that recent trading history largely drives activity predictions while asset selection depends more on available evidence.
By Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu
The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.
By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau
arXiv:2606. 08285v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling.
By Junyi Yao, Zihao Zheng
arXiv:2601.15322v3 Announce Type: replace-cross
Abstract: Tool-using agents can repeat a final decision while changing their recorded execution. We introduce the Determinism-Faithfulness Assurance Ha...
By Raffi Khatchadourian
arXiv:2607. 20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time.
By Raffi Khatchadourian
arXiv:2608.29372v1 Announce Type: new
Abstract: Retrospective backtests provide a limited test of adaptive trading agents: they cannot rule out historical contamination, expose sensitivity to a singl...
By Xiangxin Luo, Chengtian Hong, Haohua Li, Yongyi Xie
arXiv:2606. 02528v1 Announce Type: cross Abstract: Large language models now power robo-advisors and trading agents, yet whether they carry built-in biases toward specific assets is largely untested.
By Wenbin Wu
arXiv:2606. 31461v1 Announce Type: new Abstract: Niche asset markets, such as Counter-Strike 2 (CS2) weapon skins, are small, volatile, and heavily driven by community discussions and platform rules.
By Yao Shi, Kingfung Luo, Nan Tang, Yuyu Luo
arXiv:2606. 31522v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment.
By Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie
arXiv:2607. 03386v1 Announce Type: new Abstract: Agentic AI systems are increasingly used to edit, refine, and repair decision policies, but evaluating these edits is difficult when per-state expert action labels are unavailable.
By Peiying Zhu, Sidi Chang
arXiv:2608. 14825v1 Announce Type: cross Abstract: Frontier LLM agents increasingly transact on behalf of separate principals, often using natural language rather than structured APIs.
By Zeyuan Li (Massachusetts Institute of Technology), Lukas Petersson (Andon Labs), Alessandro Acquisti (Massachusetts Institute of Technology), Michiel A. Bakker (Massachusetts Institute of Technology)
arXiv:2606. 30449v1 Announce Type: new Abstract: Probes on model internals could help monitor agentic systems if they identify harmful text or tool actions before those actions are generated.
By Max Fomin, Elad David, Amit LeVi