arXiv Machine Learning

DeFiFlowBench: Benchmarking and Improving Safe Executability in Natural-Language DeFi Workflow Synthesis

DeFiFlowBench is a benchmark comprising 207 natural‑language prompts for synthesizing DeFi workflows, evaluating graph coverage, configuration completeness, and declared safety predicates, and testing trade configurations on a local EVM. The study shows that direct, constrained, and few‑shot prompting still yield unsafe executions, and that a slippage bound derived from a quote does not prevent price impact. The proposed Koan‑Safe system—combining a prompt‑only intent parser, a replaceable generator, and structural repair—achieves a higher static safety proxy score and records no unsafe executions on the benchmark, while ablation studies reveal the limits of default safety thresholds and the need for explicit trade protections. whyItMatters":"The results demonstrate that current prompting methods can still authorize costly trades and that explicit safety mechanisms like Koan‑Safe are necessary to prevent unsafe DeFi workflow executions."

arXiv Machine Learning
5d ago

SafeTune: A Unified Faithful Library for Auditing and Repairing Safety Drift in Fine-Tuned LLMs

SafeTune is a source‑available library that consolidates four safety‑intervention paradigms—post‑hoc weight recovery, safety‑constrained fine‑tuning, gradient‑based unlearning, and inference‑time steering—into a single, configuration‑driven workflow. It offers shared interpretability, evaluation, and deployment tools, and its modular registry allows easy addition of new methods, benchmarks, judges, models, and fine‑tuning domains. The authors demonstrate SafeTune with controlled comparisons and case studies in finance and medical deployments, showing how it characterizes safety drift, evaluates interventions on refusal‑behavior and capability metrics, and supports calibrated or layered mitigation.

By Pratinav Seth, Saisab Sadhu, Anshul Kaushal, Vinay Kumar Sankarapu
arXiv AI
Aug 19

PACE: Policy-Attested Contract Execution for Safe AI Agents in Decentralized Finance

The paper introduces PACE (Policy‑Attested Contract Execution), a framework that sits between large‑language‑model (LLM) based autonomous AI agents and on‑chain DeFi operations. PACE defines typed transaction intents, a deterministic policy verifier, and signed Policy Decision Records (PDRs) that cryptographically bind an approved intent, policy, and simulation report to the exact on‑chain execution bytes, providing replay and expiration protection. In evaluations across 40 tasks and six baselines, PACE achieves zero unsafe executions and zero false positives, outperforming unguarded agents by a large margin.

By Rabimba Karanjai (Larry), Yang Lu (Larry), Richard Williamson (Larry), Hemanth Hm (Larry), Prakhar Mehrotra (Larry), Lei Xu (Larry), Weidong (Larry), Shi
arXiv AI
Sep 17

Compositional Policy Violations: When Step-Level Compliance Fails In Agentic AI Workflows

The paper introduces the concept of Compositional Policy Violations (CPVs), where each step in an agentic AI workflow passes its individual compliance check, yet the overall execution violates higher‑level policies such as referral thresholds or authority limits. It categorizes CPVs into four types—Authority Creep, Threshold Laundering, Cumulative Sum Violation, and Context Collapse—and argues that the appropriate remedy depends on where the guarded quantity changes. To address this, the authors propose a provenance‑aware runtime architecture that evaluates policies over complete execution traces, recomputing guarded quantities from raw provenance rather than relying on step‑level outputs.

By Ashwini Kurady, Sri Sai Charith Grandhi, Rajesh Gupta, Sumit Mamoria