arXiv AI

Formal Verification of Agentic Systems over Operational Data

arXiv:2608. 03609v1 Announce Type: new Abstract: Agentic systems driven by large language models (LLMs) are increasingly deployed in real-world workflows where they act on persistent operational data.

arXiv AI
Sep 2

Towards Agentic Cloud Engineering: Graph and Loop Engineering with a Zero-Trust Agent Harness

The paper introduces Agentic Cloud Workflow Engineering, a framework that converts natural‑language agentic cloud‑engineering tasks into validated code repositories and verified cloud deployments. It separates graph engineering for long‑horizon workflow progression, loop engineering for bounded diagnosis and recovery, and agent harness engineering for zero‑trust execution. Experiments on Google Cloud show that executions either produce a verified deployment or an auditable terminal failure within bounded recovery limits.

By Sagar Srinivas Sakhinana, Venkataramana Runkana
arXiv AI
Aug 18

Specifying AI-SDLC Processes: A Protocol Language for Human-Agent Boundaries

arXiv:2606. 20615v3 Announce Type: replace Abstract: AI agents now act as first-class members of the software development lifecycle, but the instruments teams use to direct them enforce nothing: process encoded in prompts is flexible but unenforceable, while workflow formalisms are enforceable but do not model autonomous agents.

By Ylli Prifti, Pasquale De Meo, Alessandro Provetti
arXiv AI
Sep 17

Symbolic Temporal Supervision of LLM Agents Using Contracts

ContrAgent is a contract‑based framework that provides symbolic temporal supervision for large language model agents. It records an agent’s tool‑call sequence as a trace of checkable predicates and formalizes desired behaviors with assume‑guarantee contracts expressed in linear temporal logic over finite traces (LTLf). Each contract is compiled into a deterministic finite automaton that both gates actions online and evaluates recorded traces offline, enabling deterministic, reproducible verdicts and significantly lower per‑call latency compared to existing LLM‑judge and rule‑based guardrail baselines.

By Yifeng Xiao, Pierluigi Nuzzo