arXiv AI By Jiajun Jiang, Sharon Zheng, Natan Vidra, Spurthi Setty

Beyond Pass@k: Measuring Reliability and Security of Agentic Code Generation

Read the original on arXiv AI →

arXiv:2608. 14711v1 Announce Type: new Abstract: AI coding agent benchmarks rank agents with the Chen et al.

Summary generated by The Flow from the publisher's feed. The full article lives at arXiv AI.