arXiv AI By Yinsheng Yao, Hongxiang Zhang, Weixi Tong, Tianyi Zhang

FLARE: Fine-Grained Diagnostic Feedback for LLM Code Refinement

Read the original on arXiv AI →

arXiv:2606. 03852v1 Announce Type: cross Abstract: Large language models often generate code with bugs.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.

Hugging Face Trending Papers
Jul 24

Evaluating and Mitigating the Misguidance Effect of Buggy Code in LLM-Generated Unit Tests

While Large Language Models (LLMs) show great promise for automating unit test generation, recent studies suggest that the quality of generated tests can be negatively impacted when models are prompted with buggy code. This paper presents a new metric to quantitatively measure the "misguidance effect," a phenomenon where buggy code steers LLMs toward generating tests that validate its erroneous behavior rather than expose it.