arXiv AI

FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

FDE-Bench is a benchmark that tests large language model agents on 136 deployment‑configuration tasks involving Docker, Compose, and Kubernetes, in both greenfield and diagnose‑and‑repair scenarios. Agents submit declarative artifacts that are rebuilt and redeployed in a clean environment, and four binary check layers evaluate build, readiness, behavior, and specification conformance without an LLM judge. The benchmark includes a release gate, detailed check annotations, adversarial strategies, and reports that state‑of‑the‑art models resolve 52.9–75.0 % of tasks, while zero‑intelligence baselines solve none.

arXiv AI
2d ago

From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Praxa is an evidence‑bound harness for governed AI agent execution that explicitly represents states such as proposal, authority, dispatch, verified external effect, and promotion through deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion. The authors report four evidence lanes: a repository‑local audit passing all unit and Workerd tests; a pilot on 12 curated tasks where both baseline and reliability‑layer arms passed 17 of 36 trials; a coordination‑proxy comparison where both baseline and a source‑authored candidate completed all 180 trials with equal accuracy but the candidate used fewer tokens and steps; and deployed source/configuration evidence showing bounded reflection, recall accounting, memory compilation, and tool‑health paths. None of the evidence demonstrates superiority in security, safety, or user benefit. whyItMatters:"Praxa provides a testable architecture that makes authority‑to‑effect transitions explicit, offering a framework for verifying AI agent behavior, though current evidence does not prove improved security or performance."

By Stefan G. Creadore
arXiv AI
Jul 1

An Executable Benchmarking Suite for Tool-Using Agents

arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.

By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang
arXiv AI
6d ago

ScopeBench: Do Agents Preserve Engagement Boundaries Under Goal Pressure?

ScopeBench is a new benchmark comprising 30 dead‑end agentic security tasks designed to test whether autonomous agents respect engagement boundaries under goal pressure. Each task is presented twice: once without a scope to measure raw hacking capability, and once with a natural‑language scope to assess scope adherence. The benchmark uses deterministic verification for scopeless runs and a calibrated agentic judge for scoped runs, revealing significant gaps between capability and adherence across eight tested models.

By Shane Caldwell, Max Harley, Ads Dawson, Michael Kouremetis, Vincent Abruzzo, Will Pearce
Hugging Face Trending Papers
Aug 18

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It first checks functional claims against evidence and evaluates artifacts against nine safety properties, then tests admitted skills in a controlled environment to capture execution traces and identify failures. The approach achieves perfect precision and recall in vulnerability detection, significantly reduces attack success rates, and boosts task effectiveness and security rates in skill generation benchmarks.

arXiv AI
Sep 11

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

AgentAudit is an open, extensible framework that evaluates the full lifecycle of AI agents, assessing planning, tool selection, execution, memory, and reasoning across ten dimensions such as instruction integrity, security, and alignment. Unlike existing benchmarks that focus on single aspects, AgentAudit analyzes the entire execution trace to attribute failures to specific stages. The framework was tested on five large language models, revealing significant differences in trustworthiness even among models with similar task‑completion performance.

By Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar