Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking
arXiv:2606. 29088v1 Announce Type: cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models.
arXiv:2606. 09956v1 Announce Type: cross Abstract: The rapid adoption of LLM-powered code generation has dramatically accelerated software development, yet effective verification methods remain severely underdeveloped.
arXiv:2606. 29088v1 Announce Type: cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models.
arXiv:2609.38402v1 Announce Type: cross Abstract: We introduce C2C (From Codebase to Culprit), a framework for precise bug localization that progressively reduces the debugging search space across mu...
arXiv:2605. 17965v2 Announce Type: replace-cross Abstract: Bug localization remains a key bottleneck for large language model (LLM)-based software maintenance, where accurately identifying faulty code is essential for debugging, root cause analysis, triage, and automated program repair (APR).
arXiv:2606. 03852v1 Announce Type: cross Abstract: Large language models often generate code with bugs.
arXiv:2606. 23877v1 Announce Type: cross Abstract: Jupyter Notebooks are an increasingly popular coding environment used across many domains, especially in Python-based data science and scientific computing.
arXiv:2507. 22580v2 Announce Type: replace-cross Abstract: Automated Program Repair (APR) seeks to automatically correct software bugs without requiring human intervention.
arXiv:2608. 14065v1 Announce Type: cross Abstract: Background: Software bugs remain a critical challenge in development, necessitating effective Automated Program Repair (APR) techniques.
arXiv:2604. 11950v2 Announce Type: replace-cross Abstract: While recent LLM-based agents can identify many candidate bugs in source code, their reports remain static hypotheses that require manual validation, limiting the practicality of automated bug detection.
arXiv:2607. 00990v1 Announce Type: cross Abstract: Large language model (LLM)-based software engineering agents are increasingly developed to resolve software issues by generating patches from issue reports and code repositories.
arXiv:2512. 22827v2 Announce Type: replace-cross Abstract: Code often suffers from performance bugs.
The paper introduces SWE-Flux, a repository‑level benchmark designed to test large language models’ ability to reason about runtime behavior. It contains 480 execution‑grounded instances from 12 real Python repositories, with gold answers automatically harvested from instrumented test executions. Evaluation of five LLMs shows the task remains difficult, with the best model achieving only 37% accuracy, and the benchmark can generate challenging variants through input perturbation.
The paper investigates how large language models (LLMs) handle bug fixing versus problem solving in competitive programming. Using a dataset of ~3,000 Codeforces submissions and their human fixes, the authors compare LLM-generated patches to human patches and assess whether LLMs prefer to modify buggy code or generate new solutions. Results show that LLMs often alter more lines than necessary and sometimes produce entirely new solutions, performing better when allowed to generate solutions from scratch rather than patching existing code.