arXiv AI

PAWS: Policy-driven Agentic World Simulation

PAWS is a new dataset for policy-driven agentic world simulation that covers 36 verified U.S. financial and economic policy episodes. It includes 12,727 policy-linked news records and 65,291 stakeholder actions, each linked to supporting news and represented by a multi-layer event frame with interaction mode, financial-action family, semantic attributes, and taxonomic mappings. The dataset aligns actions with daily market-return context and has been validated by AI and human reviewers, demonstrating high agreement on interaction mode and revealing challenges in detecting rare stakeholder actions.

arXiv Machine Learning
Sep 11

AI Economist Agent: An Agentic Framework for Evidence-Based Economic and Financial Analysis with RAG, Knowledge Graphs, and Large Language Models

The paper introduces an AI economist agent that integrates large language models, retrieval‑augmented generation, knowledge graphs, and quantitative models to conduct evidence‑based economic and financial scenario analysis. The framework orchestrates LLM agents to plan analyses, retrieve relevant evidence, and structure economic mechanisms, while registered quantitative models produce numerical outcomes and predefined tests validate intermediate results for inclusion in the final report. Applied to European macro‑financial stress scenarios and bank capital analysis, the empirical study demonstrates the agent’s ability to combine flexible evidence retrieval and scenario construction while maintaining traceability to sources and explicit model calculations.

By Masahiro Kato
arXiv Machine Learning
Sep 14

FINESSE: An Agent-Based Simulator and Benchmark Dataset for Multimodal Financial Event Sequences

FINESSE is an agent‑based simulation framework that generates synthetic, structured datasets of multiple interdependent financial event streams, such as transactions, payments, account status changes, and policy interventions. Each stream has its own action space, schema, and variable types, and the streams are coupled through agents’ evolving latent states, allowing temporally rich interactions. The accompanying FINESSE‑Bench dataset supports four tasks—balance forecasting, transaction fraud detection, missed payment prediction, and next event prediction—and baseline results are provided using various time‑series and event‑sequence methods.

By Tyler Farnan, Benjamin Eng, Adam Abate, Xirui Hou, Rizal Fathony, Nam H. Nguyen, Senthil Kumar
arXiv AI
Jun 12

Fin-RATE: A Real-world Financial Analytics and Tracking Evaluation Benchmark for LLMs on SEC Filings

arXiv:2602. 07294v4 Announce Type: replace-cross Abstract: With the increasing deployment of Large Language Models (LLMs) in the finance domain, LLMs are increasingly expected to parse complex regulatory disclosures.

By Yidong Jiang, Junrong Chen, Eftychia Makri, Jialin Chen, Peiwen Li, Ali Maatouk, Leandros Tassiulas, Eliot Brenner, Bing Xiang, Rex Ying
arXiv Computation and Language
Aug 27

FinRiskAtlas: Decision-Aligned Evaluation of Large Language Models for Financial Risk Review

FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.

By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang