arXiv AI

Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation

arXiv:2608. 14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al.

arXiv AI
Sep 11

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

AgentAudit is an open, extensible framework that evaluates the full lifecycle of AI agents, assessing planning, tool selection, execution, memory, and reasoning across ten dimensions such as instruction integrity, security, and alignment. Unlike existing benchmarks that focus on single aspects, AgentAudit analyzes the entire execution trace to attribute failures to specific stages. The framework was tested on five large language models, revealing significant differences in trustworthiness even among models with similar task‑completion performance.

By Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar
arXiv AI
Sep 17

A Study of the Reliability of Agentic AI-Generated Programs

The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.

By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
arXiv AI
2d ago

From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Praxa is an evidence‑bound harness for governed AI agent execution that explicitly represents states such as proposal, authority, dispatch, verified external effect, and promotion through deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion. The authors report four evidence lanes: a repository‑local audit passing all unit and Workerd tests; a pilot on 12 curated tasks where both baseline and reliability‑layer arms passed 17 of 36 trials; a coordination‑proxy comparison where both baseline and a source‑authored candidate completed all 180 trials with equal accuracy but the candidate used fewer tokens and steps; and deployed source/configuration evidence showing bounded reflection, recall accounting, memory compilation, and tool‑health paths. None of the evidence demonstrates superiority in security, safety, or user benefit. whyItMatters:"Praxa provides a testable architecture that makes authority‑to‑effect transitions explicit, offering a framework for verifying AI agent behavior, though current evidence does not prove improved security or performance."

By Stefan G. Creadore
arXiv AI
6d ago

ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

ScopeBench is a new benchmark comprising 30 dead‑end agentic security tasks designed to test whether autonomous agents respect engagement boundaries under goal pressure. Each task is presented twice: once without a scope to measure raw hacking capability, and once with a natural‑language scope to assess scope adherence. The benchmark uses deterministic verification for scopeless runs and a calibrated agentic judge for scoped runs, revealing significant gaps between capability and adherence across eight tested models.

By Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce