arXiv AI

Towards Assurance Closure in AI-Native Large-Scale Agile Software Development

arXiv:2608. 07317v1 Announce Type: cross Abstract: The AI-Native Manifesto envisions large-scale agile software development in which humans increasingly govern intent, risk, and exceptions while agents execute more of the engineering process.

arXiv AI
Aug 24

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

The paper introduces Spec-Driven Agentic Development (SDAD), a framework that leverages large language models to ingest extensive functional requirement documents and repository context in a single workflow, turning specification quality into the engine for autonomous software delivery. SDAD blends disciplined upfront formalisation with rapid implementation, encompassing intent capture, machine‑readable specifications, agentic synthesis, and multi‑agent verification with human sign‑off. It positions AI‑code as a fourth production paradigm, compares it to traditional Waterfall and Agile approaches, and extends the model to team role evolution, quantitative governance metrics, and a staged migration blueprint for practical adoption.

By Vu Hung Nguyen, Thanh Nguyen
arXiv AI
Sep 24

From Agent Output to Authorized Transition

The paper introduces the Agile‑V Assurance Spine, a cross‑domain transition contract designed to manage the assurance of outputs from agentic engineering systems across software, firmware, and PCB domains. It specifies that evidence is only accepted when it demonstrates required properties through an authoritative source profile, is tightly bound to the exact artifact and policy baseline, stays current with declared dependencies, and meets risk‑appropriate independence and authority. Gate decisions are recorded as receipts, approvals and exceptions are scope‑ and time‑bounded, and authorization is rechecked at the effect boundary before any merge, deployment, flashing, release, or fabrication step.

By Christopher Koch
arXiv AI
2d ago

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.

By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)
arXiv AI
Jun 9

Semantic Quorum Assurance: Collective Certification for Non-Deterministic AI Infrastructure

arXiv:2606. 08021v1 Announce Type: cross Abstract: As large language model (LLM) agents are integrated into autonomous cloud operations, distributed systems face a semantic reliability problem: proposer agents can generate production mutations, such as modifying IAM policies, opening firewall security groups, or executing data exports, that are syntactically valid and statically authorized but operationally unsafe.

By Jun He, Deying Yu
arXiv AI
Aug 19

Runtime Governance for Agentic AI: Action-Boundary Control with Trusted Provenance and Fail-Closed Execution

The paper introduces Aegis, a runtime governance system for agentic AI that treats model outputs as action proposals and mediates them through a trusted decision layer before tool execution. Aegis evaluates proposals against active policy, resolves provenance server‑side, fails closed under uncertainty, and routes selected cases through a Senate‑style settlement process. In a sandbox evaluation across 6,300 rows, Aegis prevented all governed mock‑tool applications and risky side‑effect completions, preserving provenance and quorum evidence for all settled cases.

By Adam Mazzocchetti
arXiv AI
Sep 25

Graph, Loop, and Harness Engineering for Zero-Trust Agentic Data Engineering and Analytical Processing

The paper introduces two zero‑trust frameworks for cloud data engineering and analytical processing. The first, Zero‑Trust Agentic Data Engineering, automatically generates, deploys, and verifies complete data‑engineering solutions from natural‑language tasks, requiring evidence from repositories, deployments, runtimes, and policies. The second, Zero‑Trust Agentic OLAP, combines governed data preparation with verified online analytical processing, allowing production promotion only after rigorous validation and evidence‑bound approval, and ensuring analytical outputs are released only after same‑snapshot execution, exact result equivalence, deterministic grounding, and reflection. Both frameworks rely on three core abstractions—graph engineering for evidence‑gated workflow structure, loop engineering for bounded recovery, and agent‑harness engineering for zero‑trust execution—and are evaluated under nominal execution, controlled failures, bounded recovery, and policy‑constrained conditions to measure verified completion, recovery, authorization enforcement, production promotion, and verified OLAP execution.

By Sagar Srinivas Sakhinana, Venkataramana Runkana
arXiv AI
Jun 8

EvoClaw: Evaluating AI Agents on Continuous Software Evolution

arXiv:2603. 13428v2 Announce Type: replace-cross Abstract: With AI agents increasingly deployed as long-running systems, it becomes essential to autonomously construct and continuously evolve customized software to enable interaction within dynamic environments.

By Gangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan, Yuhong Liu, Yuxin Yang, Dhruv Parikh, Rajgopal Kannan, Le Cong, Mengdi Wang, Qian Zhang, Viktor Prasanna, Xiangru Tang, Xingyao Wang