arXiv:2603.10807v2 Announce Type: replace-cross
Abstract: Existing LLM safety evaluations rely on binary attack-success rates and domain-agnostic taxonomies, leaving regulated Banking, Financial Serv...
By Fabrizio Dimino, Bhaskarjit Sarmah, Stefano Pasquali
FinRiskAtlas is a Chinese-language benchmark designed to evaluate large language models (LLMs) for financial risk review by focusing on decision‑aligned tasks rather than generic financial knowledge. It contains 9,742 instances across 53 task families, including 42 domain‑knowledge families and 11 downstream review operations defined by explicit evaluation contracts. The extended FinRisk‑Ask framework replays 680 pre‑action states from 104 professional trajectories, withholding future evidence during inference to assess evidence‑state control and request targeting. Results across 33 model configurations show that operation‑level evaluation yields distinct rankings and that knowledge‑based shortlisting can incur significant regret, while frequent use of the Ask branch does not necessarily improve evidence acquisition, highlighting gaps in broad financial capability scores.
By Suyang Zhong, Jingzhe Zhu, Qi Xu, Liyao Sun, Yin Wang, Qingqing Sun, Shuai Chen, Tianyi Zhang
Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate because financial tabular data involve domain-speci...
arXiv:2607. 04103v1 Announce Type: cross Abstract: The release of SR 26-2 marks a significant modernization of U.
By Yiqing Wang, Yixin Kang, Luyun Lin, Siqi Mao
arXiv:2608.24551v1 Announce Type: cross
Abstract: Machine learning models are widely used in financial fraud and credit-risk detection, yet their adversarial robustness remains difficult to evaluate...
By Xitong Zeng, Zhaoge Bi, Yitian Yang, Huaming Chen, Quan Z. Sheng
arXiv:2607. 04103v3 Announce Type: replace-cross Abstract: Generative artificial intelligence is moving from general-purpose experimentation toward specialized applications across banking, capital markets, insurance, payments, and wealth management.
By Dennis Mao, Alessandra Lin, Yixin Kang, Yiqing Wang
arXiv:2608. 14329v1 Announce Type: cross Abstract: Principle-based regulation, with evaluative standards such as "fair, clear, and not misleading" or "deliver good outcomes", cannot be reduced to binary predicates, and LLM-as-judge is increasingly used as the substitute.
By Dipankar Sarkar
arXiv:2608. 04077v1 Announce Type: new Abstract: Evaluating financial AI agents requires criteria aligned with real professional work.
By Ben Wang, Kang Zhou, Lifan Guo, Feng Chen, Chi Zhang
arXiv:2608. 16386v1 Announce Type: cross Abstract: Financial agents must do more than recall domain knowledge: they must be both reliable, executing precise operations over grounded evidence, and executive, sustaining long-horizon research whose conclusions remain auditable.
By Agent Team, B. Zhang, Yaze Geng, Lei Tang, Yaoyang Yi, Zonghan Wu, Yifan Hu, Kun Wang, Qingsong Wen, Yilei Shao
arXiv:2607. 27853v2 Announce Type: replace-cross Abstract: Powered by advances in LLMs and autonomous agents, deep research has become one of the most widely adopted agentic products.
By Yijia Xiao, Rujun Han, Yanfei Chen, Zifeng Wang, Ke Jiang, Zhongying CuiZhu, Vishy Tirumalashetty, Wei Wang, Burak Gokturk, Tomas Pfister, Chen-Yu Lee
arXiv:2609.36474v1 Announce Type: cross
Abstract: In regulated industries like consumer finance, seemingly harmless user queries can exploit large language model vulnerabilities, triggering safety fa...
By Rikhiya Ghosh, Himanshu Kumar, Sriram Venkatapathy, Sahil Wadhwa, Alexandre G. R. Day, Pranab Mohanty
arXiv:2605. 23955v3 Announce Type: replace Abstract: Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility.
By Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du, Yi Zheng, Shouxi Ren, Deepayan Chakrabarti, Dengdu Jiang