arXiv AI

Fin-Analyst at FinMMEval 2026 Task 3: A Live Hybrid Trading Agent with LLM Specialists and Rule-Based Signals

arXiv:2607. 12233v1 Announce Type: cross Abstract: Large language model (LLM) trading agents show promising performance in equity markets, yet remain narrowly focused on US equities with little evidence from live deployment.

arXiv AI
Jul 1

FinPersona-Bench: A Benchmark for Longitudinal Psychometric Stability of Autonomous Financial Agents

arXiv:2606. 31522v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly deployed as autonomous financial agents initialized with explicit behavioral mandates such as "preserve capital" or "avoid speculative bets" that are meant to govern every decision throughout deployment.

By Muhammad Usman Safder (Steve), Ayesha Gull (Steve), Rania Elbadry (Steve), Fan Zhang (Steve), Yankai Chen (Steve), Xueqing Peng (Steve), Xue (Steve), Liu, Preslav Nakov, Zhuohan Xie
arXiv Machine Learning
Aug 13

TradingMoE: Routing the Right Experts in Evolving Markets

arXiv:2608. 11785v1 Announce Type: new Abstract: Large language models (LLMs) have shown strong potential for financial analysis and trading, but direct trading remains challenging because the predictive capabilities required can vary across assets, decision fields, and market conditions.

By Chang Zhou, Xingtong Yu, Minbin Huang, Zhennan Wu, Yuan Fang, Hong Cheng, Xinming Zhang
arXiv AI
Jun 9

Beyond Agent Architecture: Execution Assumptions and Reproducibility in LLM-Based Trading Systems

arXiv:2606. 08285v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling.

By Junyi Yao, Zihao Zheng
arXiv AI
Sep 1

Can LLMs Take the Pulse of the Economy? A Real-Time Evaluation of LLM Nowcasts on Macroeconomic Indicators

The paper introduces LiveMacroEval, a live benchmark that tests large language model (LLM) agents’ ability to produce hourly nowcasts for sixteen major U.S. macroeconomic indicators before their official release. It evaluates LLM performance against institutional nowcasts, Bloomberg ECOS consensus, and an auto-ARIMA baseline using a LiveMacro Score linked to announcement-window equity returns and a LiveBetting Score from simulated Polymarket-style trading. Over six months, state-of-the-art LLMs with web search achieved overall accuracy comparable to professional benchmarks, though performance varied across indicators.

By Xinyue Zhao, Ruiyi Zhang, Liqin Ye, Rui Cao, Pengtao Xie, Sudheer Chava
arXiv AI
Jul 14

Can Agentic Trading Systems Pay for Their Own Intelligence?

arXiv:2607. 10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value.

By Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
arXiv AI
Sep 10

What LLM Trading Agents Actually Do in Production: A Six-Month, Population-Scale Record from Two Fleets

The paper reports a six‑month, population‑scale measurement of autonomous language‑model trading agents operating in two production fleets: DX Terminal Pro, with 3,505 user‑funded vaults trading real ETH in Base memecoin markets, and the DXAP live alpha fleet, with 500–599 user‑created agents trading Hyperliquid perpetuals. Across roughly 7.5 million single‑model invocations and 231,638 multi‑tool turns, the study finds that operating layer design, risk sliders, and leaderboard boundaries drive behavior more than strategy text; agents are volatility‑blind in sizing, capture little upside, and show no directional edge compared to a retail benchmark. The analysis includes regression discontinuity, permutation nulls, and a 17‑rule methodology canon to validate the findings.

By T. J. Barton, Chris Constantakis, Patti Hauseman, Annie Mous, Alaska Hoffman, Brian Bergeron, Hunter Goodreau