arXiv AI By Cristina Improta, Pietro Liguori, Domenico Cotroneo

What is the Difference Between Me and You? Benchmarking the Quality Gap Between Human-Written and AI-Generated Code

Read the original on arXiv AI →

The Flow has not summarised this story yet — read it at arXiv AI.

arXiv Computation and Language
Sep 25

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how large language models (LLMs) handle bug fixing versus problem solving in competitive programming. Using a dataset of ~3,000 Codeforces submissions and their human fixes, the authors compare LLM-generated patches to human patches and assess whether LLMs prefer to modify buggy code or generate new solutions. Results show that LLMs often alter more lines than necessary and sometimes produce entirely new solutions, performing better when allowed to generate solutions from scratch rather than patching existing code.

By Alexandru Stefan Stoica, Traian Rebedea, Marian Cristian Mihaescu
Hugging Face Trending Papers
Sep 24

Large Language Models for Programming: Actually Fixing or Reimplementing Incorrect Code?

The paper investigates how Large Language Models (LLMs) handle bug fixing compared to human-written patches by analyzing about 3,000 Codeforces submissions. It finds that LLMs often modify more lines than necessary and sometimes produce entirely new solutions, and that they solve more problems correctly when generating solutions from scratch rather than patching existing code. The study highlights implications for AI‑assisted programming tools, suggesting a shift toward incremental problem‑solving strategies.

arXiv AI
Sep 25

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

The paper introduces a dataset comprising the complete development history of a 21,000-line Python tool created solely by Claude AI, without any human-authored code or tests. It also presents two code‑provenance tracing tools, three taxonomies for instruction intent, commit provenance, and response reliability, and applies these to analyze the dataset. Findings include that CLI instructions differ from IDE‑chat instructions, development is largely proactive, 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite, and about 1 in 4–5 interactive responses contain factual errors.

By Douglas Leith