arXiv AI By Jiajie He, Jiangyuan Hong, Xintong Chen, Dongling Ni, Wenjin Liu

Are LLMs Good Financial User Simulators? Multi-view Investor Logic Alignment (MILA)

Read the original on arXiv AI →

The study examines whether large language models (LLMs) can accurately simulate individual financial users by conducting a longitudinal paper‑trading experiment with 80 participants. Using a rolling next‑day prediction protocol, the researchers compared LLM predictions to a simple recent‑activity persistence baseline across multiple behavioral fidelity levels, from trade occurrence to asset selection and portfolio outcomes. Results show that no LLM consistently outperforms the baseline, with fidelity decreasing at finer behavioral granularity, and that recent trading history largely drives activity predictions while asset selection depends more on available evidence.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 15

Are LLMs Good Financial User Simulators? A Preliminary Study

Large language models (LLMs) are being tested as simulators of individual financial decision-making. In a controlled paper‑trading experiment with 120 volunteers, the study evaluated whether an LLM could predict a participant’s next‑day trading action, chosen security, and transaction size using only pre‑cutoff information. Results showed that including market context improved predictions of actions and tickers, but sizing remained challenging, and the models tended to over‑predict hold actions, under‑predict sells, and simplify multi‑security trades.

By Jiajie He, Jiangyuan Hong, Dongling Ni, Wenjin Liu, Xintong Chen
arXiv AI
6d ago

The Price of Thought: Does Test-Time Reasoning Pay in LLM Trading?

The study investigates whether adding inference-time reasoning to large language models (LLMs) improves trading performance. Using a controlled experiment across DeepSeek, GPT, and Gemini models, the authors varied reasoning effort while keeping other variables constant and evaluated over a full year of U.S. equities under three input conditions. Results show that additional reasoning does not reliably increase net portfolio returns and can even lead to nonmonotonic performance and unstable outcomes.

By Jiayi Chen, Guiling Wang