arXiv:2603.16365v3 Announce Type: replace
Abstract: We study alpha factor mining, the automated discovery of predictive signals from noisy, non-stationary market data-under a practical requirement th...
By Qinhong Lin, Ruitao Feng, Yinglun Feng, Zhenxin Huang, Yukun Chen, Zhongliang Yang, Linna Zhou, Binjie Fei, Jiaqi Liu, Yu Li
arXiv:2608.30192v1 Announce Type: new
Abstract: Traditional finance relies on experts to hand-craft factors through a principled process grounded in economic rationale. Recent LLM-based multi-agent s...
By Hyeonjin Kim, Minseok Kim, Seunghyeon Jung, Sujin Pyo, Huisu Jang, Woojin Lee
arXiv:2607. 26642v1 Announce Type: new Abstract: Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery.
By Jingyang Yi, Jian Yang, Yifei Jin, Yuqi Li, Jian Li
Automated alpha mining has increasingly adopted large language model (LLM) agents for factor generation and iterative discovery. However, existing LLM-based systems often delegate both factor construction and search decisions to the agent itself, without an explicit exploration space or a principled mechanism for navigating that space.
Alpha‑R1 introduces a reinforcement‑learning aligned large language model framework that performs context‑aware alpha screening by semantically gating candidate factors against a dynamic market state description. The model, trained with group relative policy optimization using realized portfolio returns as reward, selects a sparse subset of factors whose economic rationale matches current market conditions. In a 12‑month out‑of‑sample test, Alpha‑R1 achieved annualized returns of 47.87% on the S&P 500 and 40.57% on the CSI 300, with Sharpe ratios of 1.62 and 2.23, demonstrating the effectiveness of semantic factor reranking in non‑stationary markets.
By Zuoyou Jiang, Li Zhao, Rui Sun, Ruohan Sun, Zhongjian Li, Jing Li, Daxin Jiang, Zuo Bai, Cheng Hua
AlphaDiverse is a framework that enhances large language model–based multi‑agent systems for alpha factor mining by addressing cost, availability, and confidentiality constraints. It generates diverse research paths through varied environments and post‑training local agents, then fine‑tunes these agents with supervised learning and optimizes them jointly using a GRPO method that balances predictive quality and diversity. The approach limits research feedback to inner‑period data and evaluates a frozen model on outer‑period data to avoid test‑set tuning, demonstrating competitive prediction and broader exploration across four Chinese stock universes.
By Qingzhuo Wang, Zikun Wei, Zhihua Wei, Wen Shen
arXiv:2606. 12843v2 Announce Type: replace Abstract: We present an interpretable machine learning pipeline to decompose cross-sectional equity return predictability into auditable factor contributions.
By Xiao Han, Yao Xiao, Zhen Zhang, Moxuan Zheng
We present an interpretable machine learning pipeline to decompose Cross-Sectional Equity Return Predictability into auditable factor contribution. We apply an XGBoost model with TreeSHAP attribution and conduct stress testing on 3632 Chinese A-share stocks from 2009 until 2019.
arXiv:2508. 13174v2 Announce Type: replace Abstract: Formula alpha mining, which generates predictive signals from financial data, is critical for quantitative investment.
By Hongjun Ding, Binqi Chen, Jinsheng Huang, Taian Guo, Zhengyang Mao, Guoyi Shao, Lutong Zou, Luchen Liu, Ming Zhang
arXiv:2608. 12841v1 Announce Type: cross Abstract: We study recursive self-improvement at the level of quantitative-investment research: whether an autonomous system can use evidence from earlier experiments to improve the hypotheses and candidates proposed in later iterations.
By Jiacheng Guo, Suozhi Huang, Yunlong Gao, Zihao Li, Jian Ge, Xu Kuang, Mengdi Wang
The paper introduces Agentic Empirical Asset Pricing (AEAP), a framework where autonomous LLM agents conduct the entire scientific discovery process for asset pricing. It outlines AEAP’s core components, critiques current evaluation methods that only test outputs, and proposes a new reference architecture with rigorous standards for factor discovery and out‑of‑sample backtesting. Using this framework, the authors evaluate SEADS against five baselines on US equity panels, finding no single metric consistently ranks the systems and highlighting the need for multi‑axis evaluation and rolling re‑execution to assess reliability of the discovery process.
By Yingjian Pan, Xiaowei Ding, Kay Giesecke
arXiv:2608. 06618v1 Announce Type: cross Abstract: Current portfolio construction methods are either agnostic to the effects of idiosyncratic shocks (standard factor models) or to the latent data structure driving systematic returns (recent graph-based approaches).
By Sara Chehab, Giorgos Iacovides, Parisa Yazdanparast, Danilo Mandic