arXiv Machine Learning

On the Relation between Code Quality and Machine Learning Performance: A Large-scale Empirical Study

The study examined 265,363 Kaggle notebooks to explore how code quality relates to machine learning performance. Using Pylint and SonarQube, it found that general Python code quality shows negligible correlation with performance, while ML‑specific violations have a small negative association with performance. Popularity and author expertise do not predict code quality or performance, though competition expertise correlates with better performance and fewer ML‑specific violations.

arXiv AI
Sep 25

Between the Commits: Process, Error, and Claim Reliability in a Wholly AI-Authored Codebase

The paper introduces a dataset comprising the complete development history of a 21,000-line Python tool created solely by Claude AI, without any human-authored code or tests. It also presents two code‑provenance tracing tools, three taxonomies for instruction intent, commit provenance, and response reliability, and applies these to analyze the dataset. Findings include that CLI instructions differ from IDE‑chat instructions, development is largely proactive, 14.3% of AI code‑generation events contain errors later caught by the AI‑authored test suite, and about 1 in 4–5 interactive responses contain factual errors.

By Douglas Leith
arXiv AI
Jun 30

SWE-fficiency: Can Language Models Optimize Real-World Repositories on Real Workloads?

arXiv:2511. 06090v3 Announce Type: replace-cross Abstract: Optimizing the performance of large-scale software repositories demands expertise in code reasoning and software engineering (SWE) to reduce runtime while preserving program correctness.

By Jeffrey Jian Ma, Milad Hashemi, Amir Yazdanbakhsh, Kevin Swersky, Ofir Press, Enhui Li, Vijay Janapa Reddi, Parthasarathy Ranganathan