arXiv AI By Weihang Ding, Junfei Zhan, Yueting Li, Qirong Guo

FDE-Bench: Evaluating LLM Agents for Deployment Environment Configuration

Read the original on arXiv AI →

FDE-Bench is a benchmark that tests large language model agents on 136 deployment‑configuration tasks involving Docker, Compose, and Kubernetes, in both greenfield and diagnose‑and‑repair scenarios. Agents submit declarative artifacts that are rebuilt and redeployed in a clean environment, and four binary check layers evaluate build, readiness, behavior, and specification conformance without an LLM judge. The benchmark includes a release gate, detailed check annotations, adversarial strategies, and reports that state‑of‑the‑art models resolve 52.9–75.0 % of tasks, while zero‑intelligence baselines solve none.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
2d ago

From Proposal to Verified Effect: Praxa, an Evidence-Bound Harness for Governed AI Agent Execution

Praxa is an evidence‑bound harness for governed AI agent execution that explicitly represents states such as proposal, authority, dispatch, verified external effect, and promotion through deterministic admission, brokered execution, external read‑back, reconciliation, and reviewed promotion. The authors report four evidence lanes: a repository‑local audit passing all unit and Workerd tests; a pilot on 12 curated tasks where both baseline and reliability‑layer arms passed 17 of 36 trials; a coordination‑proxy comparison where both baseline and a source‑authored candidate completed all 180 trials with equal accuracy but the candidate used fewer tokens and steps; and deployed source/configuration evidence showing bounded reflection, recall accounting, memory compilation, and tool‑health paths. None of the evidence demonstrates superiority in security, safety, or user benefit. whyItMatters:"Praxa provides a testable architecture that makes authority‑to‑effect transitions explicit, offering a framework for verifying AI agent behavior, though current evidence does not prove improved security or performance."

By Stefan G. Creadore
arXiv AI
Jul 1

An Executable Benchmarking Suite for Tool-Using Agents

arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.

By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv AI
Aug 19

TRUSS: Towards Task-Reliable and User-Safe Automated Agent Skill Generation

TRUSS is a framework that generates and verifies automated agent skills, ensuring they are both functionally effective and safe. It evaluates candidate skills against source evidence and nine safety properties, then tests them in a controlled environment to capture execution traces and identify failures. The system iteratively refines skills based on these results, achieving high precision in vulnerability detection and significantly improving task performance and security rates.

By Zhibo Zhang, Zhen Ouyang, Ling Shi, Kailong Wang