arXiv AI

Diffs vs. Whole Files: An Empirical Comparison of Iterative Edit-Based and Direct Generation for Flutter/Dart Code Models

The paper compares two output regimes for code-editing language models: direct generation, where the model outputs the entire modified file, and iterative diff-based generation, where the model emits a sequence of localized edits. Experiments on Flutter/Dart tasks show that direct generation consistently outperforms diff-based generation across metrics such as compilation success, token efficiency, and quality judgments. However, diff-based generation can be competitive for short, spatially localized edits, particularly in refactoring and error-handling tasks with few edit steps.

arXiv AI
Sep 4

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

The paper investigates the problem of over‑editing by large language models when repairing code, showing that even state‑of‑the‑art models like GPT‑5.5 frequently rewrite more code than necessary. Using a benchmark of 400 BigCodeBench problems with controlled AST corruptions, the authors quantify excess edits and demonstrate that a simple preservation instruction can reduce unnecessary changes and improve pass rates. They further explore training strategies, finding that reinforcement learning yields the best balance between edit fidelity and performance retention, highlighting edit fidelity as a distinct, measurable dimension of code‑repair quality.

By Tongyao Zhu, Wei Hern Lim, Min-Yen Kan
Hugging Face Trending Papers
Sep 3

When Models Edit Too Much: On the Fidelity of Minimal Code Edits

The paper investigates the problem of over‑editing by large language models (LLMs) when repairing code, showing that even state‑of‑the‑art models like GPT‑5.5 often rewrite more code than necessary. Using a benchmark of 400 BigCodeBench problems with controlled AST‑level corruptions, the authors quantify over‑editing and demonstrate that a simple preservation instruction can significantly reduce excess edits and cognitive complexity while improving Pass@1. They further explore post‑training strategies, finding that reinforcement learning yields the best balance between edit fidelity and performance retention, thereby establishing edit fidelity as a distinct, measurable dimension of code‑repair quality.

arXiv Computation and Language
Sep 4

CROCODIL: Cross-Model Code Editing with LLMs

CROCODIL is a post‑training framework designed to improve cross‑model code editing with large language models (LLMs). It addresses the problem that different LLMs, trained on distinct datasets, often make excessive edits when applied to code generated by another model. By combining a similarity reward that discourages large changes with an execution reward that ensures build and test success, CROCODIL encourages smaller, functionally correct edits.

By Linghan Zhong, Aditya Thimmaiah, Jayanth Srinivasa, Milos Gligoric, Junyi Jessy Li
arXiv Machine Learning
1d ago

When Edit Flows are Edit Jumps: replicating Edit Flows and EvoFlows

arXiv:2609.18745v1 Announce Type: new Abstract: Antibody lead optimization calls for a small, bounded set of edits to an existing candidate: substitutions, but also insertions and deletions. Edit-bas...

By Gabriel B\'en\'edict, Melanie Buechler, Gerard Riera-Sol\`a, Chlo\'e de Ancos, Yves Gaetan Nana Teukam, Moritz Freidank
arXiv Computation and Language
4d ago

Local Edits, Global Ripples: Replay-Informed Policy Adaptation for Workflow Synthesis

The paper introduces RIPPLE, a method for adapting workflow-synthesizing agents through prompt-policy editing without retraining the underlying model. RIPPLE diagnoses failed execution trajectories, maps failures to specific policy segments, and restricts edits to those segments. It then evaluates candidate edits in isolation and replays only those that remain safe after composition, achieving up to a 23.1% improvement in validation success on a synthetic benchmark and positive gains on additional language‑model backbones.

By Manqing Mao, Hong Wang, Samson Koelle, Jie Yuan, Zhuoer Wang, James Feng, Yanjun Lin, Daniel Edmiston, Nikki Lijing Kuang, Zhecheng Sheng, Wei Niu