arXiv AI By Lucas Jing, Xinqi Wang, Liao Zhang, Simon S. Du

PBT-Bench: Benchmarking AI Agents on Property-Based Testing

Read the original on arXiv AI →

arXiv:2605. 15229v3 Announce Type: replace-cross Abstract: Existing code benchmarks measure whether an agent can produce any test that reproduces a known bug, or whether it can produce a patch that fixes a described issue.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.