Trustworthy Runtime Error Healing in Real-World Repositories: A Benchmark and Guardrail
Read the original on arXiv AI →The paper introduces HealBench, a benchmark of 265 runtime errors from 18 real-world repositories, and HealGuard, a safety framework that restricts healing code to an analyzable subset of Python and applies static and dynamic taint analysis. Using these tools, the authors evaluate a dedicated healing method and three general coding agents powered by different LLM backbones, achieving a 38.11% resume rate and a 28.68% test‑pass rate, while HealGuard flags 17.4% of successful healings as potentially unsafe. The study demonstrates that current LLM agents can meaningfully repair real repository crashes, but also highlights significant safety concerns that the Guardrail framework can detect, albeit with a high false‑positive rate.
Machine-generated by The Flow from the publisher's headline and feed description — not written or checked by a human. The full article lives at arXiv AI.