arXiv AI

An AI Security Agent for Banking: Multi-Vector Fraud and AML Detection Across Retail and Corporate Accounts

arXiv:2606. 17555v1 Announce Type: cross Abstract: Banks simultaneously face signature-based fraud (card-not-present attacks, account takeover, ATM cloning) and behavioural financial crime (structuring, layering, mule networks, business email compromise) -- two threat families with fundamentally different detection requirements.

arXiv AI
4d ago

Agentic Commerce Bench: Measuring Fraud Detection for Agents That Spend Money

The paper introduces the Agentic Commerce Bench (ACB), a benchmark for measuring fraud in AI agents that autonomously spend money. It presents a taxonomy of agentic commerce fraud, a dataset of twenty fraud classes derived from real production data, and an open‑source detector stack called gordonguard for auditing and replaying hostile counterparties. The study shows that current reasoning layers and security scanners perform poorly on many classes, highlighting the need for better detection mechanisms.

By Ankit Srivastava, Debjyoti Paul
arXiv AI
Aug 20

FraudBench: Stress-Testing Policy-Grounded Banking Agents Against Adaptive Fraud

FraudBench is a new benchmark that tests policy‑grounded banking conversational agents against adaptive fraud scenarios. It uses a dual‑control framework and a 698‑document internal policy corpus, presenting 150 adversarial scenarios (107 public, 43 held‑out) that require agents to manage mutable account state and tool access while preventing identity, authorization, and trust manipulation. Preliminary results on four agents show attack‑security rates between 49% and 65%, highlighting weaknesses in money‑mule and first‑party fraud detection.

By Dheeraj Mohandas Pai, Lu Xian
arXiv AI
Sep 11

Cyber-Financial Contagion: Modeling the Propagation of an AI Vendor Compromise Through the Banking System

The paper examines how a breach of a single AI vendor—used by banks for fraud screening, credit decisions, AML triage, customer analytics, and internal support—can spread through operational, informational, and financial links, ultimately causing losses that resemble a traditional banking crisis. It introduces a four‑layer heterogeneous network linking AI vendors, banks, interbank exposures, and customer accounts, and presents CFC‑Prop, a stochastic epidemic‑and‑clearing model that reproduces heavy‑tailed loss distributions and sensitivity to patch latency on a synthetic dataset of 60 vendors, 220 banks, and 1,400 interbank exposures. Additionally, the authors develop CFC‑GNN, an early‑warning graph‑based model that predicts high‑cascade‑risk vendors with AUROC 0.82 and AUPRC 0.60, and they release all code, data, and scripts for reproducibility.

By Alex Leytes
arXiv AI
Sep 2

A Hybrid Insider Threat Detection Framework Combining Multi-Agent Simulation, Layered SIEM Correlation, and Theory-of-Mind Reasoning

The paper introduces a hybrid insider threat detection framework that combines multi-agent simulation, layered SIEM correlation, trust‑adaptive thresholds, behavioral and communication forensics, and Theory‑of‑Mind reasoning. It evaluates four variants—Layered SIEM‑Core, Cognitive‑Enriched SIEM, Evidence‑Gated SIEM, and an Enron‑calibrated version—showing progressively higher actor‑level F1 scores and reduced false positives, especially with evidence gating. Domain‑shift tests reveal that an Enron‑trained email classifier does not transfer to other domains, but fine‑tuning achieves high F1, and scalability tests confirm stable performance up to 1,000 agents.

By Firdous Kausar, Asmah Muallem, Naw Safrin Sattar, Mohamed Zakaria Kurdi