arXiv AI By Douglas Leith

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

Read the original on arXiv AI →

The paper introduces a dataset comprising the complete development history of a 21,000-line Python tool created solely by Claude AI, without any human-authored code or tests. It also presents two code‑provenance tracing tools, three taxonomies for instruction intent, commit provenance, and response reliability, and applies these to analyze the dataset. Findings include that CLI instructions differ from IDE‑chat instructions, development is largely proactive, 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite, and about 1 in 4–5 interactive responses contain factual errors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.

Hugging Face Trending Papers
Sep 24

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

The paper introduces a dataset of the complete development history of a 21,000-line Python tool built entirely by Claude AI, accompanied by two code‑provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability. Analysis reveals that user CLI instructions differ from IDE‑chat instructions, focusing more on comprehension, planning, and consultation; code development is largely proactive; 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite; and roughly one in four to five of the AI’s interactive responses contain factual errors.