The paper introduces an AI economist agent that integrates large language models, retrieval‑augmented generation, knowledge graphs, and quantitative models to conduct evidence‑based economic and financial scenario analysis. The framework orchestrates LLM agents to plan analyses, retrieve relevant evidence, and structure economic mechanisms, while registered quantitative models produce numerical outcomes and predefined tests validate intermediate results for inclusion in the final report. Applied to European macro‑financial stress scenarios and bank capital analysis, the empirical study demonstrates the agent’s ability to combine flexible evidence retrieval and scenario construction while maintaining traceability to sources and explicit model calculations.
By Masahiro Kato
arXiv:2609. 03553v1 Announce Type: new Abstract: Policy analysis requires more than predicting whether a proposal will pass: it requires identifying who will be affected, how those actors respond, and what follows.
By Linh Le, Melanie Bui, My Chiffon Nguyen, Zachary Schlosser, David Williams-King
FINESSE is an agent‑based simulation framework that generates synthetic, structured datasets of multiple interdependent financial event streams, such as transactions, payments, account status changes, and policy interventions. Each stream has its own action space, schema, and variable types, and the streams are coupled through agents’ evolving latent states, allowing temporally rich interactions. The accompanying FINESSE‑Bench dataset supports four tasks—balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction—and baseline results are provided using various time‑series and event‑sequence methods.
By Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar
arXiv:2607. 26588v1 Announce Type: new Abstract: The rapid development of large language models (LLMs) has renewed interest in agent-based modeling (ABM).
By Shaopeng Wei, Yufei Cheng, Wenxi Sun, Yepeng Ding, Yu Zhao, Gang Kou
arXiv:2607. 11141v1 Announce Type: new Abstract: Large language models (LLMs) based agents are beginning to participate in portfolio construction and market analysis, where decisions must be justified under evolving information and risk constraints.
By Changlun Li, Peixian Ma, Qiqi Duan, Zhenyu Lin, Peineng Wu
arXiv:2607. 19409v1 Announce Type: new Abstract: Recent advances in large language models have accelerated deployment of agentic systems in operational finance.
By Wolfgang M. Pauli, Sarah Panda, Kidus Admassu, Said Bleik, Ademola Okerinde, Jeremy Reynolds
arXiv:2607. 20491v1 Announce Type: new Abstract: Standard evaluation benchmarks measure what a tool-using agent decides, not whether it arrives at that decision through the same process each time.
By Raffi Khatchadourian
arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.
By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
arXiv:2507.19364v3 Announce Type: replace
Abstract: The integration of Large Language Models (LLMs) into social simulation has generated considerable enthusiasm, but also raises substantial methodolo...
By Patrick Taillandier, Jean Daniel Zucker, Arnaud Grignard, Benoit Gaudou, Nghi Quang Huynh, Haojia Kong, Alexis Drogoul
arXiv:2609.34211v2 Announce Type: replace
Abstract: In financial fraud detection, rich semantic context can provide important evidence for transaction behavior modeling and fraud reasoning. However,...
By Linbo Shao, Huilin He, Yating Lou, Dawei Cheng
arXiv:2609.23703v1 Announce Type: cross
Abstract: Financial language models can transform unstructured firm-specific news into structured decision signals, but financial AI research lacks an integrat...
By Kemal Kirtac
FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang