arXiv:2608. 11683v1 Announce Type: new Abstract: AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow.
By Yuhao Zhang, O. Ozan Koyluoglu, Thejas Venkatesh, Richard Diehl Martinez, Vishank Bhatia, Arash Alidoust, Ashwin Paranjape
arXiv:2607. 11141v1 Announce Type: new Abstract: Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints.
By Changlun Li, Peixian Ma, Qiqi Duan, Zhenyu Lin, Peineng Wu
AI agents are increasingly deployed for professional investment research, yet no benchmark captures the complexity of the full investor workflow. Existing benchmarks mainly target financial data extraction, a narrow slice that current models have largely saturated, while reference-based metrics and generic LLM-as-a-judge scoring fall short on the open-ended, long-form answers that real analyst queries demand.
arXiv:2603.16365v3 Announce Type: replace
Abstract: We study alpha factor mining, the automated discovery of predictive signals from noisy, non-stationary market data-under a practical requirement th...
By Qinhong Lin, Ruitao Feng, Yinglun Feng, Zhenxin Huang, Yukun Chen, Zhongliang Yang, Linna Zhou, Binjie Fei, Jiaqi Liu, Yu Li
arXiv:2607. 20645v1 Announce Type: cross Abstract: We introduce Frontier Financial Judgement, a challenging new benchmark developed in collaboration with professional equity analysts to assess agents' ability to replicate expert human judgements.
By Joshua Harris
EvoTS-Agent is a self‑evolving large language model agent designed for autonomous change‑point detection in financial time series. It begins with curated exploratory data analysis to set up candidate models, then iteratively refines detection pipelines using three operators—Revision, Alternative Strategy, and Recombination—guided by validation feedback. Across four benchmark datasets, EvoTS-Agent consistently outperforms existing LLM‑based agents and achieves a 100% execution success rate with all tested backbone LLMs.
EvoTS-Agent is a self‑evolving large language model agent designed for autonomous change‑point detection in financial time series. It begins with curated exploratory data analysis to set up candidate models, then iteratively refines its detection pipeline using three operators—Revision, Alternative Strategy, and Recombination—guided by validation feedback. Across four benchmark datasets, EvoTS-Agent consistently outperforms existing LLM‑based agents and achieves a 100% execution success rate on all tested backbone LLMs.
By Lei Jiang, Ye Wei, Xinyu Xi, Jordan Langham-Lopez, Yifan Bao, Raad Khraishi, Yihao Ang, Anthony K. H. Tung, Lukasz Szpruch, Hao Ni
arXiv:2606. 26346v1 Announce Type: new Abstract: Agentic benchmarks have emerged across general-purpose and domain-specific settings, including finance, coding, law, and drug discovery, yet energy-domain evaluations remain largely limited to static knowledge recall.
By David Akinpelu, Akintonde Abbas, Rereloluwa Alimi, Ayodeji Lana
arXiv:2605. 05580v2 Announce Type: replace Abstract: Quantitative trading agents have demonstrated substantial promise in automating factor discovery, signal aggregation, and portfolio execution.
By Yishuo Yuan, Jiayi Sheng, Sirui Zeng, Jiaqi Wang, Jiaheng Liu
arXiv:2608.30192v1 Announce Type: new
Abstract: Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent s...
By Hyeonjin Kim, Minseok Kim, Seunghyeon Jung, Sujin Pyo, Huisu Jang, Woojin Lee
arXiv:2606. 11851v1 Announce Type: new Abstract: Open-ended scientific discovery asks agents to move beyond executing analyses for predefined questions.
By Jiayao Chen, Shi Liu, Linyi Yang
arXiv:2608. 12841v1 Announce Type: cross Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations.
By Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang