The paper introduces a dataset comprising the complete development history of a 21,000-line Python tool created solely by Claude AI, without any human-authored code or tests. It also presents two code‑provenance tracing tools, three taxonomies for instruction intent, commit provenance, and response reliability, and applies these to analyze the dataset. Findings include that CLI instructions differ from IDE‑chat instructions, development is largely proactive, 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite, and about 1 in 4–5 interactive responses contain factual errors.
By Douglas Leith
arXiv:2605.29442v2 Announce Type: replace-cross
Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectori...
By Ningzhi Tang, Chaoran Chen, Gelei Xu, Yiyu Shi, Yu Huang, Collin McMillan, Tao Dong, Toby Jia-Jun Li
arXiv:2604.20779v2 Announce Type: replace
Abstract: AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful...
By Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo
arXiv:2606. 12329v1 Announce Type: new Abstract: AI coding assistants now support a growing share of software work, from quick scripts to production applications.
By Ripon Chandra Malo, Tong Qiu
arXiv:2607. 21832v1 Announce Type: cross Abstract: Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows.
By Iren Mazloomzadeh, Mohammad Mehdi Morovati, Foutse Khomh
arXiv:2608. 05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review.
By Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
arXiv:2609.12708v2 Announce Type: replace-cross
Abstract: AI coding assistants are becoming co-authors of production software, yet their evaluation centers on functional correctness, leaving open whe...
By Cristina Improta, Pietro Liguori, Domenico Cotroneo
The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.
By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.
By Guanqun Yang, Wei Yang, Xueqing Liu
arXiv:2607. 29516v1 Announce Type: cross Abstract: AI coding agents are generating code at volumes that exceed the capacity of traditional peer review.
By Chandra Maddila, Mashrur Rashik, Euna Mehnaz Khan, Smriti Jha, James Saindon, Nachi Nagappan, Peter C. Rigby
arXiv:2606. 18168v1 Announce Type: cross Abstract: Software practitioners increasingly use AI coding agents that generate test code alongside production code in open source pull requests (PRs).
By Dipayan Banik, Kowshik Chowdhury, Shazibul Islam Shamim
arXiv:2609.05677v1 Announce Type: cross
Abstract: Lifelong LLM agents increasingly rely on external skill artifacts as one element for preserving and reusing capabilities over time. These skills (usu...
By Chen Shen, Estevam Hruschka