The paper introduces the Agentic Commerce Bench (ACB), a benchmark for measuring fraud in AI agents that autonomously spend money. It presents a taxonomy of agentic commerce fraud, a dataset of twenty fraud classes derived from real production data, and an open‑source detector stack called gordonguard for auditing and replaying hostile counterparties. The study shows that current reasoning layers and security scanners perform poorly on many classes, highlighting the need for better detection mechanisms.
By Ankit Srivastava, Debjyoti Paul
arXiv:2606. 25007v1 Announce Type: new Abstract: Financial fraud detection in digital banking requires reasoning over multiple heterogeneous event streams -- transactions, login sessions, risk signals -- that individually appear benign but collectively reveal fraudulent patterns.
By Mohammadamin Dashti Moghaddam, Nick Sciarrilli
arXiv:2607. 19350v1 Announce Type: new Abstract: Financial institutions face significant challenges in detecting sophisticated money laundering patterns, such as smurfing and layering, due to extreme data imbalance (0.
By Mariam Zakaria Moussa Ali
FraudBench is a new benchmark that tests policy‑grounded banking conversational agents against adaptive fraud scenarios. It uses a dual‑control framework and a 698‑document internal policy corpus, presenting 150 adversarial scenarios (107 public, 43 held‑out) that require agents to manage mutable account state and tool access while preventing identity, authorization, and trust manipulation. Preliminary results on four agents show attack‑security rates between 49% and 65%, highlighting weaknesses in money‑mule and first‑party fraud detection.
By Dheeraj Mohandas Pai, Lu Xian
arXiv:2607. 13469v1 Announce Type: cross Abstract: The banking sector increasingly relies on automated systems to monitor electronic transactions for signs of fraud, yet conventional rule-based approaches struggle with high false-positive rates and offer no justification for their outputs, limiting their utility for compliance teams.
By Anupa Lodhi
arXiv:2607. 19266v1 Announce Type: cross Abstract: Fraud detection systems must scale with rising transaction volume while remaining explainable and reviewable.
By Rahil Sharma
The paper examines how a breach of a single AI vendor—used by banks for fraud screening, credit decisions, AML triage, customer analytics, and internal support—can spread through operational, informational, and financial links, ultimately causing losses that resemble a traditional banking crisis. It introduces a four‑layer heterogeneous network linking AI vendors, banks, interbank exposures, and customer accounts, and presents CFC‑Prop, a stochastic epidemic‑and‑clearing model that reproduces heavy‑tailed loss distributions and sensitivity to patch latency on a synthetic dataset of 60 vendors, 220 banks, and 1,400 interbank exposures. Additionally, the authors develop CFC‑GNN, an early‑warning graph‑based model that predicts high‑cascade‑risk vendors with AUROC 0.82 and AUPRC 0.60, and they release all code, data, and scripts for reproducibility.
By Alex Leytes
arXiv:2608. 15447v1 Announce Type: new Abstract: Mobile money has widened financial access across Sub-Saharan Africa and enlarged the surface for money-laundering and terrorism-financing (ML/TF) activity in ecosystems dominated by high-volume, low-value transactions.
By Emmanuel Nahimana, Ya\'e Ulrich Gaba
arXiv:2506.11635v2 Announce Type: replace-cross
Abstract: Credit card fraud mitigation plays a significant role in modern society. While fraud detection systems are essential, they often struggle to...
By Shaun Shuster, Eyal Zloof, Asaf Shabtai, Rami Puzis
arXiv:2604. 17420v2 Announce Type: replace-cross Abstract: Money laundering poses severe risks to global financial systems, driving the widespread adoption of machine learning for transaction monitoring.
By Keyang Chen, Mingxuan Jiang, Yongsheng Zhao, Zeping Li, Zaiyuan Chen, Weiqi Luo, Zhixin Li, Sen Liu, Yinan Jing, Guangnan Ye, Xihong Wu, Hongfeng Chai
The paper introduces a hybrid insider threat detection framework that combines multi-agent simulation, layered SIEM correlation, trust‑adaptive thresholds, behavioral and communication forensics, and Theory‑of‑Mind reasoning. It evaluates four variants—Layered SIEM‑Core, Cognitive‑Enriched SIEM, Evidence‑Gated SIEM, and an Enron‑calibrated version—showing progressively higher actor‑level F1 scores and reduced false positives, especially with evidence gating. Domain‑shift tests reveal that an Enron‑trained email classifier does not transfer to other domains, but fine‑tuning achieves high F1, and scalability tests confirm stable performance up to 1,000 agents.
By Firdous Kausar, Asmah Muallem, Naw Safrin Sattar, Mohamed Zakaria Kurdi
arXiv:2609.14234v1 Announce Type: cross
Abstract: Financial fraud in corporate transaction networks has grown more coordinated and harder to detect with rule-based engines and with classical learning...
By Sergei, Komarov