The paper introduces a dataset of the complete development history of a 21,000-line Python tool built entirely by Claude AI, accompanied by two code‑provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability. Analysis reveals that user CLI instructions differ from IDE‑chat instructions, focusing more on comprehension, planning, and consultation; code development is largely proactive; 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite; and roughly one in four to five of the AI’s interactive responses contain factual errors.
arXiv:2605.29442v2 Announce Type: replace-cross
Abstract: AI coding agents increasingly act directly within software environments, yet existing analyses of their failures rely on benchmark trajectori...
By Ningzhi Tang, Chaoran Chen, Gelei Xu, Yiyu Shi, Yu Huang, Collin McMillan, Tao Dong, Toby Jia-Jun Li
arXiv:2606. 12329v1 Announce Type: new Abstract: AI coding assistants now support a growing share of software work, from quick scripts to production applications.
By Ripon Chandra Malo, Tong Qiu
arXiv:2604.20779v2 Announce Type: replace
Abstract: AI coding agents are being adopted at scale, yet we lack empirical evidence on how people actually use them and how much of their output is useful...
By Joachim Baumann, Vishakh Padmakumar, Xiang Li, John Yang, Diyi Yang, Sanmi Koyejo
arXiv:2607. 21832v1 Announce Type: cross Abstract: Recent advances in large language models and their rapid adoption across software engineering tasks have made Artificial Intelligence (AI) coding agents an integral component of modern software development workflows.
By Iren Mazloomzadeh, Mohammad Mehdi Morovati, Foutse Khomh
arXiv:2609.12708v2 Announce Type: replace-cross
Abstract: AI coding assistants are becoming co-authors of production software, yet their evaluation centers on functional correctness, leaving open whe...
By Cristina Improta, Pietro Liguori, Domenico Cotroneo
arXiv:2608. 05179v1 Announce Type: cross Abstract: Large language model (LLM) agents are increasingly used across the scientific research lifecycle: ideation, literature search, experiment design and execution, analysis, manuscript drafting, and review.
By Tianyu Ding, Aditya Nannapaneni, Bingfan Liu, Ling Zhang
The paper investigates the reliability of software produced by agentic AI by comparing AI-generated versions of ten well-known Linux utilities to their human-written counterparts. Using fuzz testing (both black-box and coverage-guided AFL++), the authors find that AI-generated code is often as reliable or more reliable than the latest human versions, with fewer memory errors but a higher incidence of hangs. The study emphasizes that robust AI-generated software requires careful prompting, skilled human oversight, and that the AI workflow can serve as a cost-effective specification for sustainable code.
By Ayesha Shafique, Barton P. MIller, Elisa R. Heymann
arXiv:2606. 18168v1 Announce Type: cross Abstract: Software practitioners increasingly use AI coding agents that generate test code alongside production code in open source pull requests (PRs).
By Dipayan Banik, Kowshik Chowdhury, Shazibul Islam Shamim
WatchPoint is a simulated‑user system that generates and runs diagnostic scripts against a live web application, producing structured observations to guide coding model retries. Unlike prior methods that rely on screenshots or non‑executable metrics, WatchPoint operates on Web‑Bench—a benchmark of 50 multi‑file web projects with 1,000 sequential tasks verified by deterministic end‑to‑end tests. It recovers 57.6% of diagnosed tasks, matching a human tester’s 54.5% recovery rate, and identifies when such simulated feedback is beneficial or should be withheld.
By Guanqun Yang, Wei Yang, Xueqing Liu
The paper investigates whether large language models can reliably reproduce official Eurostat statistics by generating executable code. It evaluates a coding agent across four experimental conditions—task only, task plus metadata, metadata with a repair loop using execution feedback, and metadata with a retry budget but no diagnostics—using 30 natural‑language tasks spanning seven domains and datasets. Results show that success depends on semantic validation against frozen specifications, a fully specified output contract, and a retry budget, rather than on execution diagnostics alone.
By Sabina-Cristiana Necula
arXiv:2606.21804v2 Announce Type: replace-cross
Abstract: Maintainability is a core dimension of software engineering, shaping how code is written, reviewed, and developed over time. While coding age...
By Shaswat Patel, Betty Li Hou, Arun Purohit, Kai Xu, Jane Pan, He He, Valerie Chen