The paper introduces Spec-Driven Agentic Development (SDAD), a framework that leverages large language models to ingest extensive functional requirement documents and repository context in a single workflow, turning specification quality into the engine for autonomous software delivery. SDAD blends disciplined upfront formalisation with rapid implementation, encompassing intent capture, machine‑readable specifications, agentic synthesis, and multi‑agent verification with human sign‑off. It positions AI‑code as a fourth production paradigm, compares it to traditional Waterfall and Agile approaches, and extends the model to team role evolution, quantitative governance metrics, and a staged migration blueprint for practical adoption.
By Vu Hung Nguyen, Thanh Nguyen
The paper introduces the Agile‑V Assurance Spine, a cross‑domain transition contract designed to manage the assurance of outputs from agentic engineering systems across software, firmware, and PCB domains. It specifies that evidence is only accepted when it demonstrates required properties through an authoritative source profile, is tightly bound to the exact artifact and policy baseline, stays current with declared dependencies, and meets risk‑appropriate independence and authority. Gate decisions are recorded as receipts, approvals and exceptions are scope‑ and time‑bounded, and authorization is rechecked at the effect boundary before any merge, deployment, flashing, release, or fabrication step.
By Christopher Koch
arXiv:2607. 21495v1 Announce Type: new Abstract: AI agents are increasingly created inside organizations by non-engineering users through low-code, no-code, and conversational development environments.
By Natan Levy, Harel Berger
The paper introduces a deterministic AI security risk assessment framework that transforms diverse engineering artefacts into a standardized Control ID taxonomy scored on a four‑level ordinal scale. It compiles technique‑level predicates from a fixed MITRE ATLAS snapshot, linking each control to mitigation and producing traceable feasibility and impact outputs. The framework is formally verified for boundedness, totality, consistency, and monotonicity, and is evaluated on five open‑source AI projects, showing that strengthened controls lower feasibility scores while residual risks persist when core controls are missing.
By Yixuan Huang (University of Southampton, Southampton, UK), Basel Halak (University of Southampton, Southampton, UK), Boojoong Kang (University of Southampton, Southampton, UK)
arXiv:2606. 08021v1 Announce Type: cross Abstract: As large language model (LLM) agents are integrated into autonomous cloud operations, distributed systems face a semantic reliability problem: proposer agents can generate production mutations, such as modifying IAM policies, opening firewall security groups, or executing data exports, that are syntactically valid and statically authorized but operationally unsafe.
By Jun He, Deying Yu
arXiv:2609.22664v1 Announce Type: cross
Abstract: Research on large language model agents for penetration testing is evaluated almost entirely by capability: whether the agent captures a flag or repr...
By Joas Antonio dos Santos Barbosa