arXiv AI
5d ago

Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail

The paper introduces HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, and HealGuard, a safety framework that restricts healing code to an analyzable subset of Python and applies static and dynamic taint analysis. Using these tools, the authors evaluate a dedicated healing method and three general coding agents powered by different LLM backbones, achieving a 38.11% resume rate and a 28.68% test‑pass rate, while HealGuard flags 17.4% of successful healings as potentially unsafe. The study demonstrates that current LLM agents can meaningfully repair real repository crashes, but also highlights significant safety concerns that the Guardrail framework can detect, albeit with a high false‑positive rate.

By Gou Tan, Pengfei Chen, Zhensu Sun, Jieke Shi, Junkai Chen, Ting Zhang, Weifeng Sun, Junda He, Shuai Liang, Chuanfu Zhang, Lwin Khin Shar, David Lo
arXiv AI
Sep 4

PatchBench: Evaluating AI Agents for Vulnerability Patching

PatchBench introduces a benchmark to evaluate AI agents on realistic vulnerability patching tasks, addressing two key threats to validity: patch memorization and surface-level fixes that merely suppress crashes. The study finds that 25% of agent patches resemble historical developer patches, and that PoC-only validation inflates success rates by 1.83× on average. PatchBench mitigates these issues by selecting vulnerabilities whose true fixes lie outside the crash stack, migrating historical vulnerabilities into new contexts, and employing rigorous validation for security and semantic correctness.

By Chihao Shen, Jiacheng Li, Aastha Mahajan, Jeffery Siyuan Tian, Yonghwi Kwon, Yizheng Chen