arXiv:2607. 23942v1 Announce Type: new Abstract: Memory, planning, reflection, and tool use are often compared as feature labels, obscuring the control semantics that determine how an agent actually runs.
By Haodi Fan, Zucong Lan
The paper introduces PAI‑Bench, a benchmark designed to evaluate persistent AI agents on how faithfully they adhere to a versioned identity contract. It separates several dimensions—recall, composition, behavioral enactment, resistance, persistence, lineage, and role‑conditioned updates—while keeping scoring oracles independent of the target process. Experiments on synthetic profiles show that explicit cues can significantly alter the presence of identity identifiers, revealing prompt‑dependent component selection and sensitivity to startup cues.
By Zhenyu Zhao, Roy Zhao
arXiv:2608. 11632v1 Announce Type: cross Abstract: Persistent AI agents accumulate versioned state across long horizons, but storage retention alone does not identify authoritative state.
By Jun He, Deying Yu
The paper examines how agentic AI systems intended for military command and control are tested and evaluated. It reviews 240 testing practices across eight dimensions and three lifecycle stages, uncovering eight assumptions—grouped into system specifiability, stability, composability, and supervisability—whose validity is weakened by agentic properties. Consequently, test results may meet procedural standards but do not guarantee that fielded behavior matches tested behavior, leading the authors to propose ten assurance claims and suggest that uncertainty be managed through deployment‑time monitoring and defined expiry conditions.
By Ulysse Richard, Heather Frase, Sarah Cao, Di Cooke, Sebastian Kwon, Adrianna Tan
The paper introduces Runtime Assurance Contracts (RAC) as a formal policy framework for high‑risk AI agents, addressing the "assurance‑transition gap" by binding autonomy boundaries, component eligibility, evidence state, transition policy, human‑review capacity, and non‑compensatory gates. RAC allows soft metrics to influence routing while mandating retries, switches, escalations, deferrals, or stops when mandatory gates fail or are unknown, ensuring aggregate performance cannot alone authorize action. The authors define the contract, evidence record, permission rule, and five invariants, and evaluate RAC through deterministic failure‑injection studies, hand‑authored traces, and a prospective synthetic holdout, comparing it to score‑only and restricted protocol baselines.
By Serhii Zabolotnii
The paper introduces a continuous evaluation framework that assesses both outcome-level and process-level aspects of evolving enterprise AI agent skills. It applies this framework to two variants of a Business Value Determination skill, running 240 trials across multiple models, harnesses, and specifications. The results show that while most trials pass final numerical checks, a large majority still exhibit process-level deviations, and dependency attribution reduces the number of failed checks per run. The framework also provides reusable regression tests and highlights specification sensitivity across configurations.
By Ngoc Phuoc An Vo, Aarya Doshi, Vadim Sheinin
arXiv:2607. 16345v1 Announce Type: cross Abstract: Modern agentic systems increasingly rely on skills: installable packages of natural language and code that teach an LLM agent to perform a domain task.
By Tejas Singh Anand, Yuet Ying Christina Wang, Wanting Jiang, Steve Masson, Tian Zheng, Bingjie Zhou
arXiv:2607. 14275v1 Announce Type: new Abstract: Context engineering has become central to building reliable AI agents, yet it remains largely unmeasured.
By Fouad Bousetouane
arXiv:2607. 14890v1 Announce Type: new Abstract: Autonomous coding agents increasingly execute multi-step software work, but lifecycle states such as reviewed, tested, DONE, and ready-to-merge remain claims unless supported by current evidence.
By Jek Huang, Jeffery Hsia, Jiayi Sun, Freddie Shi, Wei Huang, Ian H. White
The paper introduces the Agile‑V Assurance Spine, a cross‑domain transition contract designed to manage the assurance of outputs from agentic engineering systems across software, firmware, and PCB domains. It specifies that evidence is only accepted when it demonstrates required properties through an authoritative source profile, is tightly bound to the exact artifact and policy baseline, stays current with declared dependencies, and meets risk‑appropriate independence and authority. Gate decisions are recorded as receipts, approvals and exceptions are scope‑ and time‑bounded, and authorization is rechecked at the effect boundary before any merge, deployment, flashing, release, or fabrication step.
By Christopher Koch
arXiv:2606. 30306v1 Announce Type: cross Abstract: Always-on agents are systems whose future behavior depends on durable state accumulated across earlier interactions.
By Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
arXiv:2609.16313v1 Announce Type: cross
Abstract: In agentic distributed systems, an agent may be authorized to mutate external infrastructure while lacking evidence that the mutation is ready to exe...
By Jun He, Deying Yu