arXiv AI By Thanh Luong Tuan, Abhijit Sanyal

Toward Pre-Deployment Assurance for Enterprise AI Agents: Ontology-Grounded Simulation and Trust Certification

Read the original on arXiv AI →

arXiv:2606. 04037v1 Announce Type: new Abstract: Pre-deployment verification of enterprise artificial intelligence (AI) agents remains a critical gap between large language model (LLM) capability benchmarking and production deployment.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

arXiv AI
Sep 1

Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act

The paper "Operationalising AI Regulatory Sandboxes: Activities, Requirements, and Technical Assessment under the EU AI Act" outlines a detailed framework for implementing AI Regulatory Sandboxes (AIRS) under the EU AI Act. It maps the sandbox lifecycle into 29 activities, distinguishes between a Core AIRS and an Extended AIRS that includes an AI Technical Sandbox (AITS), and derives 15 infrastructural and governance requirements linked to these activities and provider obligations. The authors also introduce the Sandbox Configurator, an open‑source tool to instantiate AITS environments, aiming to provide structured workflows for regulators, robust evaluation methods for experts, and a transparent compliance pathway for AI providers.

By Alessio Buscemi, Thibault Simonetto, Daniele Pagani, German Castignani, Maxime Cordy, Jordi Cabot
arXiv AI
Sep 3

READY or Not: Reliable Enterprise Agent Deployment

READY or Not: Reliable Enterprise Agent Deployment introduces a framework for qualifying AI agents for enterprise workflows. It measures reliability and operating cost under various oversight policies, selects the minimum‑cost policy that meets a specified reliability target, and statistically qualifies it on held‑out cases. In a clinical audit study, READY revealed that two agents with nearly identical autonomous accuracy required markedly different levels of human review to achieve the same reliability target.

By Veronica Chatrath (Christy), Bryan Zhu (Christy), Jingxuan Fan (Christy), George Pu (Christy), Soham Dinesh Tiwari (Christy), Soham Dan (Christy), Ryan Young (Christy), Yuan (Christy), Li, Yuang Yao, Apaar Shanker, Minglai Yang, Daniel Yue Zhang, Yunzhong He, Ying Liu, Chenguang Wang, Zhijun Yin, Yuan Xue