arXiv AI

Spec2COBOLRot: An Agentic-AI Degradation Loop for Realistic COBOL Corpus Generation

The paper introduces Spec2COBOLRot, an agentic AI pipeline that generates realistic COBOL programs by combining specification-driven creation with iterative degradation guided by real production code patterns and complexity targets. The authors evaluate the pipeline on three programs from different business domains, showing that it reliably produces syntactically valid code and increases structural complexity, but it does not consistently preserve business behavior. They discuss the limitations of targeting structural metrics alone and propose future work that would generate legacy programs from scratch along a simulated development history.

arXiv AI
Jun 8

EvoClaw: Evaluating AI Agents on Continuous Software Evolution

arXiv:2603. 13428v2 Announce Type: replace-cross Abstract: With AI agents increasingly deployed as long-running systems, it becomes essential to autonomously construct and continuously evolve customized software to enable interaction within dynamic environments.

By Gangda Deng, Zhaoling Chen, Zhongming Yu, Haoyang Fan, Yuhong Liu, Yuxin Yang, Dhruv Parikh, Rajgopal Kannan, Le Cong, Mengdi Wang, Qian Zhang, Viktor Prasanna, Xiangru Tang, Xingyao Wang
arXiv AI
Aug 24

SDAD: Spec-Driven Agentic Development for the AI-Native SDLC

The paper introduces Spec-Driven Agentic Development (SDAD), a framework that leverages large language models to ingest extensive functional requirement documents and repository context in a single workflow, turning specification quality into the engine for autonomous software delivery. SDAD blends disciplined upfront formalisation with rapid implementation, encompassing intent capture, machine‑readable specifications, agentic synthesis, and multi‑agent verification with human sign‑off. It positions AI‑code as a fourth production paradigm, compares it to traditional Waterfall and Agile approaches, and extends the model to team role evolution, quantitative governance metrics, and a staged migration blueprint for practical adoption.

By Vu Hung Nguyen, Thanh Nguyen
arXiv AI
5d ago

Code That Works, Environments That Don't: Measuring Environment Reproducibility in AI-Generated Software

The paper investigates how well large language models specify the software environments needed to run AI-generated code. Using a new agent protocol and a three‑layer dependency framework, the authors evaluate three coding agents across four languages and fifty tasks, finding that dependency specifications are often inconsistent, redundant, or incomplete. Agreement on dependency sets is as low as 7% for identical tasks, and newer agents show no improvement, indicating that environment specification remains a significant, unaddressed challenge in code generation.

By Bhanu Prakash Vangala, Tanu Malik
arXiv AI
Jul 7

Don't Blame the Large Language Model: How Scaffolding Evolution Shapes Coding Agent Quality

arXiv:2607. 03691v1 Announce Type: cross Abstract: Coding agents, autonomous systems that use large language models (LLMs) to resolve software engineering tasks, rely on agentic scaffolding: a middleware layer in between a developer and a large language model that orchestrates system prompts, tool execution, context management, and iterative reasoning loops.

By Oussama Ben Sghaier, Hao Li, Bram Adams, Ahmed E. Hassan
arXiv AI
Sep 11

A-JIT: Agentic Just-In-Time Software Construction

The paper introduces Agentic Just-In-Time Software Construction (A-JIT), a paradigm that replaces static software binaries with dynamic systems capable of continuous evolution. In A-JIT, an application consists of code, a runtime harness, and an embedded AI agent that observes usage and execution traces to specialize software logic, workflows, and tool interfaces for each user. This approach enables applications to dynamically generate missing implementations, create new capabilities on the fly, and adapt continuously to end‑user behavior, thereby supporting trace‑driven human‑AI co‑construction and opening a new design space for adaptive, self‑evolving software.

By Mark Marron, Earl T. Barr