arXiv AI

FLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement

arXiv:2606. 03852v1 Announce Type: cross Abstract: Large language models often generate code with bugs.

Hugging Face Trending Papers
Jul 24

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests

While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metric to quantitatively measure the "misguidance effect," a phenomenon where buggy code steers LLMs toward generating tests that validate its erroneous behavior rather than expose it.