arXiv AI

From Agent Output to Authorized Transition

The paper introduces the Agile‑V Assurance Spine, a cross‑domain transition contract designed to manage the assurance of outputs from agentic engineering systems across software, firmware, and PCB domains. It specifies that evidence is only accepted when it demonstrates required properties through an authoritative source profile, is tightly bound to the exact artifact and policy baseline, stays current with declared dependencies, and meets risk‑appropriate independence and authority. Gate decisions are recorded as receipts, approvals and exceptions are scope‑ and time‑bounded, and authorization is rechecked at the effect boundary before any merge, deployment, flashing, release, or fabrication step.

arXiv AI
Aug 24

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

The paper introduces Spec-Driven Agentic Development (SDAD), a framework that leverages large language models to ingest extensive functional requirement documents and repository context in a single workflow, turning specification quality into the engine for autonomous software delivery. SDAD blends disciplined upfront formalisation with rapid implementation, encompassing intent capture, machine‑readable specifications, agentic synthesis, and multi‑agent verification with human sign‑off. It positions AI‑code as a fourth production paradigm, compares it to traditional Waterfall and Agile approaches, and extends the model to team role evolution, quantitative governance metrics, and a staged migration blueprint for practical adoption.

By Vu Hung Nguyen, Thanh Nguyen
arXiv AI
Sep 7

Beyond Code Generation: Reliability, Verification, and Cost Economics in the Agentic Software Development Lifecycle

The paper examines how AI coding agents are evolving beyond simple autocomplete to perform complex tasks such as repository inspection, multi-file editing, tool execution, test writing, pull request creation, and long-duration work with minimal supervision. It highlights that while these agents boost coding activity, significant bottlenecks remain in review, integration, testing, security, deployment, and production operations, and that the economics of software development are shifting toward variable token, tool, sandbox, CI, and rework costs. The authors synthesize recent research and industry data to propose four engineering concepts—Agentic SDLC Throughput Paradox, Production-Qualified Change, Verification Tax, and an Agentic SDLC Control Plane—to guide the allocation of autonomy within cost, reliability, and human-attention constraints, ultimately reframing the research focus to production-qualified value per dollar, reviewer-hour, and operational risk.

By Happy Bhati
arXiv AI
2d ago

A Deterministic and Auditable AI Security Risk Assessment Framework with ATLAS Aligned Executable Rules and Formal Verification

The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.

By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)
arXiv AI
Jul 1

An Executable Benchmarking Suite for Tool-Using Agents

arXiv:2605. 11030v2 Announce Type: replace-cross Abstract: Closed-loop tool-using agents are increasingly evaluated in executable web, code, and micro-task environments, but benchmark reports often conflate workloads, action-generating drivers, and the evidence admitted for systems-facing claims.

By Zhiqing Zhong, Zhijing Ye, Jiamin Wang, Xiaodong Yu
arXiv AI
6d ago

Resource-Optimized and Energy-Aware Agentic AI Framework Anchored on Blockchain for Secure Software Supply Chains

The paper introduces a blockchain-backed agentic security framework that protects the entire software development lifecycle and the AI components monitoring it. It coordinates specialized security agents—covering source integrity, dependency and SBOM analysis, CI configuration auditing, artifact verification, and runtime policy evaluation—each powered by a large language model that interprets artifacts, reasons over tool outputs, and generates structured security reports. Every agent produces a cryptographically signed attestation recorded on a permissioned blockchain via smart contracts, creating an immutable attestation log, an agent registry, and an enforceable release‑policy module, while communication is secured through a consortium‑operated certificate authority.

By Toqeer Ali Syed, Asadullah Abdullah Khan
arXiv AI
Sep 11

Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance

The paper "Beyond Training: A Feasibility Taxonomy for Inference-Time AI Governance" presents a taxonomy of twenty inference‑time mechanisms for monitoring, verification, and enforcement, each evaluated on a four‑point readiness scale using evidence from four vendors. It applies this taxonomy to a two‑dimensional adversary model and maps the mechanisms to four governance scenarios, finding that most mechanisms are commercially available but only adequate against cooperative or low‑to‑medium‑capability users, not high‑capability state‑level deployers. The study also links inference‑stage controls to hardware‑stage mechanisms through a substitution principle and reports a second‑rater reliability of 0.74. whyItMatters":"The work identifies the current gaps and readiness of inference‑time governance tools, highlighting that existing mechanisms are insufficient against powerful adversaries and thus informing future regulatory and technical development."

By Samar Ansari