Hugging Face Trending Papers

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

Read the original on Hugging Face Trending Papers →

The paper introduces a dataset of the complete development history of a 21,000-line Python tool built entirely by Claude AI, accompanied by two code‑provenance tracing tools and three taxonomies for instruction intent, commit provenance, and response reliability. Analysis reveals that user CLI instructions differ from IDE‑chat instructions, focusing more on comprehension, planning, and consultation; code development is largely proactive; 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite; and roughly one in four to five of the AI’s interactive responses contain factual errors.

Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at Hugging Face Trending Papers.

arXiv AI
Sep 25

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

The paper introduces a dataset comprising the complete development history of a 21,000-line Python tool created solely by Claude AI, without any human-authored code or tests. It also presents two code‑provenance tracing tools, three taxonomies for instruction intent, commit provenance, and response reliability, and applies these to analyze the dataset. Findings include that CLI instructions differ from IDE‑chat instructions, development is largely proactive, 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite, and about 1 in 4–5 interactive responses contain factual errors.

By Douglas Leith