arXiv:2607. 13078v1 Announce Type: cross Abstract: LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows.
By Keyur Gabani
LLMs are now proposed for fraud detection, scam investigation, content moderation, and other trust-and-safety workflows. Much of the public literature still evaluates them as models, with less attention to their behavior as components in operational pipelines.
arXiv:2608. 08577v1 Announce Type: new Abstract: Fraud operations must allocate events among automatic approval, analyst review, and automatic blocking even though the labels needed to evaluate these actions are selective and delayed.
By Jie Deng (Tongji University, Shanghai, China)
The paper introduces the Agentic Commerce Bench (ACB), a benchmark for measuring fraud in AI agents that autonomously spend money. It presents a taxonomy of agentic commerce fraud, a dataset of twenty fraud classes derived from real production data, and an open‑source detector stack called gordonguard for auditing and replaying hostile counterparties. The study shows that current reasoning layers and security scanners perform poorly on many classes, highlighting the need for better detection mechanisms.
By Ankit Srivastava, Debjyoti Paul
The paper introduces DISCERN, a two-tier protocol for certifying that updates to production models do not increase risk. It first uses unlabeled data to detect benign updates based on disagreement rates, then selectively labels only disagreements through an anytime-valid confidence sequence. The method achieves finite-sample validity with label-complexity bounds of order ρ²/ε², demonstrating significant label savings and strong empirical performance across 14,000+ audit streams.
By Vishnu Bindu Balachandran
The paper presents a deployed system that scores blockchain addresses using their position in a massive multi‑chain transaction graph instead of relying on sanctions lists. The system operates on a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, employing a shared inductive encoder with per‑chain normalization and two scoring heads. It demonstrates label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis at a very low alert rate, and shows significant lead time over external registry events, while maintaining fast, reproducible serving performance.
By Yury Korolev
The paper audits self‑evolving financial agents—SkillOpt, Agent Workflow Memory (AWM), and ReasoningBank—by evaluating their performance, security drift, and execution‑interface mismatches in simulated e‑banking scenarios. It shows that while utility improves after evolution, exposure to malicious content and unauthorized state changes also rise, and that AWM’s text‑action envelope can disrupt tool execution, highlighting the need for comprehensive auditing beyond accuracy metrics.
By Jialong Li, Jialing Zhu
The paper presents a compliance screening system that evaluates blockchain addresses by their position in a large multi‑chain transaction graph instead of relying on sanctions lists. Using a single graph of 835 million addresses and 15.8 billion edges across five EVM chains, the system employs a shared inductive encoder with per‑chain normalization and two scoring heads, with decision thresholds set as exact quantiles of the score distribution. The authors demonstrate label‑free transfer, achieving high recall on held‑out positives for Base, Arbitrum, and Gnosis, and report significant lead‑time in flagging external registry events, efficient serving latency, and robustness checks against adversarial behavior.
arXiv:2608. 01095v1 Announce Type: new Abstract: Federated learning (FL) enables multiple intelligent devices to collaboratively train a high-accuracy model without sharing raw data.
By Hongliang Zhang, Zhongyuan Yu, Fenghua Xu, Teng Hu, Jian Meng, Jiguo Yu
The paper presents RuntimeGuard‑AI, a prototype that links each deterministic AI policy decision to its source code, writes a privacy‑minimizing record at a chosen synchronization point, and returns an Ed25519‑signed receipt indicating whether the write succeeded. After a crash, the system validates the integrity of records, manifests, shard placement, sequence continuity, and replay identity, while an independent attestation path chains committed records into signed Merkle epochs for auditor verification. Performance results on an Apple M4 Pro show high throughput (up to 27,193 requests/s) with low latency when buffering, but throughput drops and latency rises when per‑record data and full synchronization are used, illustrating a clear durability‑latency trade‑off.
By Neeraj Kumar Singh Beshane
arXiv:2607. 11607v1 Announce Type: new Abstract: Distributional reinforcement learning agents learn full return distributions that are increasingly read at face value: for interpretability, risk-sensitive control, and safety monitoring.
By Hari Prasad
The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.
By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi