arXiv AI By Manyi Wang, Junjielong Xu, Pinjia He

PAIChecker: Uncovering and Checking PR-Issue Misalignment in SWE-Bench-Like Benchmarks

Read the original on arXiv AI →

arXiv:2607. 28587v2 Announce Type: replace-cross Abstract: SWE-bench-like benchmarks are widely used for evaluating LLM's issue resolution capability.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.