arXiv AI By Vincent Schmalbach

Software Delegation Contracts: Measuring Reviewability in AI Coding-Agent Work

Read the original on arXiv AI →

arXiv:2606. 17099v1 Announce Type: cross Abstract: AI coding agents increasingly accept assigned software tasks, modify repositories under bounded authority, and return work packages for review.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
3d ago

Trust Is Not a Score: Runtime Assurance Contracts for High-Risk AI Agents

The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.

By Serhii Zabolotnii
arXiv AI
2d ago

From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Praxa is an evidence‑bound harness for governed AI agent execution that explicitly represents states such as proposal, authority, dispatch, verified external effect, and promotion through deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion. The authors report four evidence lanes: a repository‑local audit passing all unit and Workerd tests; a pilot on 12 curated tasks where both baseline and reliability‑layer arms passed 17 of 36 trials; a coordination‑proxy comparison where both baseline and a source‑authored candidate completed all 180 trials with equal accuracy but the candidate used fewer tokens and steps; and deployed source/configuration evidence showing bounded reflection, recall accounting, memory compilation, and tool‑health paths. None of the evidence demonstrates superiority in security, safety, or user benefit. whyItMatters:"Praxa provides a testable architecture that makes authority‑to‑effect transitions explicit, offering a framework for verifying AI agent behavior, though current evidence does not prove improved security or performance."

By Stefan G. Creadore
arXiv AI
Sep 25

SWE-Prometheus: Measuring Engineering Governance Improvements in Real-World Repositories

SWE-Prometheus is a new benchmark that evaluates large language model coding agents on the broader task of improving repository engineering governance, rather than just fixing individual issues. It presents fixed snapshots with open-ended objectives, requiring agents to identify risks, prioritize interventions, and verify changes across six governance dimensions. The benchmark includes 60 repositories and reports metrics such as Normalized Governance Improvement and behavior‑breakage rates, providing a nuanced view of how different models affect governance artifacts and execution‑backed improvements.

By Jiajun Wu, Leixin Sun, Zihan Tan, Yitao Liu, Shuo Li, Jiaru Qian, Shanghaoran Quan, Chuangxin Zhao, Yangxu Liao, Yang Liu, Bin Chong, Guancheng Wan
arXiv AI
Sep 18

Governance-as-Code: Translating EU AI Act Technical Requirements into Executable Compliance Pipelines for Generative AI Systems

The paper introduces Governance-as-Code (GaC), a framework that translates the EU AI Act’s technical requirements into 43 machine‑checkable acceptance criteria across six compliance modules. GaC runs within a CI/CD pipeline, producing Article‑indexed audit evidence and providing actual Rego policy code. The authors validate GaC on two enterprise deployments, showing it reproduces manual audit findings—including three penalty‑triggering violations—while reducing audit labor by about 75%.

By Rudrendu Kumar Paul, Sourav Nandy