FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
arXiv:2606. 08285v1 Announce Type: new Abstract: Large language models (LLMs) and agentic systems are increasingly proposed for financial trading, yet their reported performance remains difficult to compare because studies vary in data provenance, temporal split discipline, execution timing, turnover treatment, and transaction-cost modeling.
By Junyi Yao, Zihao Zheng
arXiv:2608. 04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work.
By Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
EvolveTrade is a self‑evolving framework that treats the system prompt of a tool‑using LLM trading agent as a text‑parameterized policy. After each update interval, a Policy Agent revises this policy using accumulated decision traces and portfolio feedback while keeping the backbone LLM fixed, allowing the agent to refine its information‑acquisition and portfolio‑construction procedures over time. Experiments across multiple market regimes and two LLM backbones show that EvolveTrade often improves Sharpe Ratio and Cumulative Return over fixed‑policy baselines, with behavioral analyses indicating increased code‑mediated analysis and regime‑relevant computations.
By Sehee Kim, Yumin Choi, Minki Kang, Sung Ju Hwang
arXiv:2606. 29771v1 Announce Type: new Abstract: LLM agents are increasingly cast as autonomous portfolio managers, and benchmarks have moved from financial question-answering to sequential trading.
By Bo Qu, Mingguang Chen
arXiv:2606. 00051v1 Announce Type: cross Abstract: Large Language Models (LLMs) are increasingly used in analytical workflows, but their suitability as exploratory data analysis (EDA) agents in business settings remains uncertain.
By Rafa{\l} {\L}ab\k{e}dzki, Patryk Miziu{\l}a, Hubert Rutkowski, Szymon Betlewski, Cezary Depta, Szymon Janowski, Jaros{\l}aw Kochanowicz, Jan Kanty Milczek
arXiv:2607. 10286v1 Announce Type: new Abstract: Large language model (LLM) agents are increasingly used in trading systems, where model reasoning, tool use, and continual decisions incur costs that are expected to produce trading value.
By Qiqi Duan, Changlun Li, Chen Wang, Fan Zhang, Mengxiang Wang, Dayi Miao, Peixian Ma, Jiangpeng Yan, Liyuan Chen, Shuoling Liu, Preslav Nakov, Yuyu Luo, Nan Tang
arXiv:2606. 23032v3 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks.
By Mostapha Benhenda
arXiv:2608. 15591v1 Announce Type: new Abstract: Large Language Model (LLM) agents deployed in production environments face a fundamental tension: the agent's behavior is frozen at deployment time, while the business rules and edge cases it must handle continue to evolve.
By Pouya Ghiasnezhad Omran, Michael Zimmermann, Duncan Cambridge, Ashmita Kapoor, Tanya Dixit
arXiv:2606. 23032v2 Announce Type: replace Abstract: Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks.
By Mostapha Benhenda
Finance Agent v2 (by Vals AI) has emerged as the reference benchmark for evaluating both Anthropic Claude and OpenAI ChatGPT frontier language models on financial tasks. However, it narrowly deals with periodic reporting from publicly traded companies (SEC 10-K and 10-Q filings), and its agentic harness relies on naive, unenriched chunk retrieval.
arXiv:2607. 19409v1 Announce Type: new Abstract: Recent advances in large language models have accelerated deployment of agentic systems in operational finance.
By Wolfgang M. Pauli, Sarah Panda, Kidus Admassu, Said Bleik, Ademola Okerinde, Jeremy Reynolds