arXiv AI

Continuous Process-Level Evaluation for Evolving Enterprise AI Agent Skills

The paper introduces a continuous evaluation framework that assesses both outcome-level and process-level aspects of evolving enterprise AI agent skills. It applies this framework to two variants of a Business Value Determination skill, running 240 trials across multiple models, harnesses, and specifications. The results show that while most trials pass final numerical checks, a large majority still exhibit process-level deviations, and dependency attribution reduces the number of failed checks per run. The framework also provides reusable regression tests and highlights specification sensitivity across configurations.

arXiv AI
Aug 24

Evaluating Skills, Not Just Agents: Agentic Continuous Evaluation of Skills

The paper introduces ACES (Agentic Continuous Evaluation of Skills), a framework that evaluates reusable skills and capability packages by running paired live trials with and without a target skill, normalizing results into the Agent Trajectory Interchange Format (ATIF), and grading six runtime metrics to compute Skill Lift. ACES demonstrates that scan-only gates miss important aspects of skill performance, while the evaluation protocol reveals significant improvements in skill execution, behavior check, and skill efficiency across 145 real skills and 947 scored cases. The open‑source NVIDIA SkillEvaluator implementation enables reproducible, repository‑native assessment of agentic skills in production environments.

By Christopher Kevin, Narendran Raghavan, Jean-Francois Puget, Roshni Malani, Meghana Puvvadi, Moshe Abramovitch, Mohit Gupta, Rama Akkiraju, Subodh Prabhu, Yogesh Dangi, Wei Luo, Seong Hee Lee
arXiv AI
Jun 2

BADGER: Bridging Agentic and Deterministic Evaluation for Generative Enterprise Reasoning

arXiv:2606. 02109v1 Announce Type: new Abstract: Enterprise AI systems that translate natural language into SQL queries and orchestrate multi-step agentic reasoning pipelines require evaluation approaches fundamentally different from academic benchmarks.

By Shannon Serrao, Soumitra Chatterjee, Dorina Strori, Abhishek Sharma, Nathan Miller
arXiv AI
Sep 11

AgentAudit: An Open, Extensible Framework for Full-Lifecycle Trust Evaluation of AI Agents

AgentAudit is an open, extensible framework that evaluates the full lifecycle of AI agents, assessing planning, tool selection, execution, memory, and reasoning across ten dimensions such as instruction integrity, security, and alignment. Unlike existing benchmarks that focus on single aspects, AgentAudit analyzes the entire execution trace to attribute failures to specific stages. The framework was tested on five large language models, revealing significant differences in trustworthiness even among models with similar task‑completion performance.

By Shrey Nag, Sachita, Abhishek Kumar Singh, Lipi Goel, Rajeshwar Singh Janwar
arXiv AI
Sep 3

READY or Not: Reliable Enterprise Agent Deployment

READY or Not: Reliable Enterprise Agent Deployment introduces a framework for qualifying AI agents for enterprise workflows. It measures reliability and operating cost under various oversight policies, selects the minimum‑cost policy that meets a specified reliability target, and statistically qualifies it on held‑out cases. In a clinical audit study, READY revealed that two agents with nearly identical autonomous accuracy required markedly different levels of human review to achieve the same reliability target.

By Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue