arXiv AI By Bal\'azs Szalontai, \'Abel Szauter, Bal\'azs M\'arton, P\'eter Verebics, Bal\'azs Pint\'er, Tibor Gregorics

Diff-Based Code Corruption using LLMs for Large-Scale Bugfix Benchmarking

Read the original on arXiv AI →

arXiv:2606. 29088v1 Announce Type: cross Abstract: There are various benchmarks to evaluate bugfixing capabilities of Large Language Models.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.