arXiv:2609.24016v1 Announce Type: new
Abstract: Commercial large language models are increasingly deployed across African fintech infrastructure for fraud detection and customer communication, yet no...
By Andrew Anogie Uduimoh, Hadiza Umar Yusuf, Oluwafemi Osho
arXiv:2607. 17797v1 Announce Type: new Abstract: Financial statements (FS) such as Balance Sheet (BS), Income Statement (IS) and Cash-flow Statement (CS) summarize the annual financial performance of a company.
By Kshitij Madhav Jadhav, Sushodhan Vaishampayan, Manoj Apte, Sachin Pawar, Nitin Ramrakhiyani, Girish Keshav Palshikar
arXiv:2607. 13469v1 Announce Type: cross Abstract: The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-based approaches struggle with high false-positive rates and offer no justification for their outputs, limiting their utility for compliance teams.
By Anupa Lodhi
arXiv:2608. 07471v1 Announce Type: cross Abstract: This study considers the task of applying artificial intelligence to recognize bank fraud.
By Bohdan Mytnyk, Oleksandr Tkachyk, Nataliya Shakhovska, Solomiia Fedushko, Yuriy Syerov
arXiv:2606. 19887v1 Announce Type: cross Abstract: Existing safety benchmarks target general adversarial scenarios but miss finance-specific risks.
By Chaeyun Kim, Daeyoung Park, Junghwan Kim, Jinyoung Jeong, Eunji Song, Yongtaek Lim, Minwoo Kim
arXiv:2607. 10317v1 Announce Type: cross Abstract: Algorithmic decision systems in financial services often rely on data proxies that inadvertently encode structural inequalities.
By Muhammad Abdullahi Said
FinRCA-Bench is a synthetic benchmark designed to evaluate evidence retrieval and reasoning in financial AI systems, specifically for accounts‑payable‑to‑bank reconciliation. It contains 2,250 cases across 14 operational tables, with 1,500 injected failures in 15 causal categories and 750 hard‑negative cases, and hides root‑cause labels and evidence contracts to isolate retrieval performance. Experiments show that retrieval architecture dramatically affects accuracy, with structured retrieval methods like Typed Provenance Graph Retrieval vastly improving macro‑recall and exact‑class accuracy compared to dense semantic retrieval or classical ML.
arXiv:2605. 23955v3 Announce Type: replace Abstract: Deploying machine learning in regulated financial environments -- credit risk, fraud detection, and anti-money laundering -- exposes critical vulnerabilities in algorithmic reproducibility.
By Ruizhe Zhou, Xiaoyang Liu, Gaoyuan Du, Yi Zheng, Shouxi Ren, Deepayan Chakrabarti, Dengdu Jiang
arXiv:2607. 04103v3 Announce Type: replace-cross Abstract: Generative artificial intelligence is moving from general-purpose experimentation toward specialized applications across banking, capital markets, insurance, payments, and wealth management.
By Dennis Mao, Alessandra Lin, Yixin Kang, Yiqing Wang
The paper presents LAAF, a Layered Accountability Architecture Framework for Large Language Model (LLM) applications, developed through a systematic review of 122 primary studies and 12 regulatory documents. It identifies five dimensions of accountability and four families of mechanisms—technical controls, human oversight, organisational governance, and documentation/traceability—each assessed for maturity. The framework is mapped onto major regulatory standards (EU AI Act, NIST AI RMF, ISO/IEC 42001) and highlights persistent gaps such as under‑specified human oversight and lack of shared accountability metrics.
arXiv:2506.11635v2 Announce Type: replace-cross
Abstract: Credit card fraud mitigation plays a significant role in modern society. While fraud detection systems are essential, they often struggle to...
By Shaun Shuster, Eyal Zloof, Asaf Shabtai, Rami Puzis
FinRCA-Bench is a deterministic synthetic benchmark comprising 2,250 accounts‑payable‑to‑bank reconciliation cases that span 14 operational tables and include 1,500 injected failures across 15 causal categories. The benchmark hides root‑cause labels and record‑level evidence contracts from models, enabling independent evaluation of evidence retrieval versus reasoning accuracy. Experiments show that retrieval architecture dramatically influences performance, with retrieval improvements raising macro‑required‑record recall from 0.83% to 77.70% and exact 16‑class accuracy from 2.05% to 72.44%.
By Pratik Ghawate